An Unverifiable World12. Contain the ConsequencesContents中文

Chapter 12: Contain the Consequences

Thesis: When you cannot prevent errors, manage their consequences. Limit the damage that something wrong and unverified can do (containment, before the fact), and make sure that if it does go wrong, you will find out (the audit trail, after the fact).

The three earlier pairs of moves were all still trying to get the thing right, even when they lowered the bar for what counts as right. This last pair admits that you cannot get it right and turns to managing failure instead. If you cannot prevent the error, shrink the damage it can do (containment, before the fact), and make sure that when it happens you can trace it (the audit trail, after the fact).

Containment: Shrink the Blast Radius

Here is the first move in its pure form. Give up on guaranteeing that the unverified thing will not fail. Instead, decide before the fact how far its failure can reach, and fence that off.

This idea is the deepest foundation of computer security. The line runs from Saltzer and Schroeder's principle of least privilege in 19751, Lampson's confinement in 19732, and Dennis and Van Horn's capability mechanism in 19667 to Denning's information flow lattice in 19765 and the security models of Bell-LaPadula3 and Biba4. The theme is the same throughout. Give a component only the minimum capability it needs to do its own job, and nothing more, so that even if it is breached or fails, it cannot do much damage. The sandbox (Goldberg et al., 1996)9, separation of duties, and defense in depth are all forms of it. In systems reliability engineering, it is the circuit breaker and the bulkhead (Nygard's Release It!15), blast-radius design, canary releases, and the error budget (Google SRE30). In finance, it is the position limit and the stop-loss. And Taleb's antifragility17 is also about capping the downside.

The unifying idea is to shift the burden from "make it not fail" (which would take the verification you do not have) to "make it survivable even when it fails." Defense in depth offers a common quantitative intuition. If $k$ layers of protection each fail independently with probability $p$, the probability that all of them fail at once is

$$p^{k},$$

which falls exponentially as you add layers. But the warning from Chapter 10 applies right away: this $p^k$ holds only when the layers fail independently. If the layers fall to the same weakness (the same kernel that gets bypassed, the same administrator password), correlation instantly reduces defense in depth to a single layer. The Fukushima nuclear accident of 2011 shows this principle at work. The plant had multiple redundancy, with main power plus backup diesel generators, but a single tsunami drowned them together. Several lines of defense that were supposed to be independent of one another fell to the same cause, and defense in depth turned out to be hollow.

Defense in depth and the Swiss cheese model: when the holes line up, failure runs straight through

This is where the move's standard failure mode lies: containment that gets bypassed. Sandboxes have escapes. Privileges quietly creep. What looks like layer upon layer of defense turns out to be a set of layers that share one hidden door. A less often mentioned failure mode is over-protection. It blocks normal work too, so people end up working around it to get things done, and security becomes a sham after all.

The Audit Trail: Make Errors Visible After the Fact

Here is the second move in its pure form. Whatever you cannot prevent, make sure you detect the moment it happens. Move the check from before the fact to after it.

Its sturdiest techniques come from cryptography. Merkle's hash tree (the Merkle tree) of 198018 and Haber and Stornetta's chained timestamps of 199119 make it impossible to alter a record quietly once it has been written down. Any change is exposed when the record is checked. Crosby and Wallach's tamper-evident log of 200923, Schneier and Kelsey's protection of logs on untrusted machines in 199820, and Bellare and Miner's forward-secure signature of 199922 make these techniques sturdier still. Certificate Transparency (RFC 6962)24 and Nakamoto's Bitcoin of 200827 are, at heart, global, append-only audit ledgers that anyone can verify. Checking whether a record is in such a tree costs only $O(\log n)$. Once again, this is the dividend of the asymmetry from Chapter 7: verifying is cheaper than producing.

The move itself is far older. Double-entry bookkeeping is one of humanity's earliest tamper-evident ledgers. In The Reckoning29, Soll argues that the ability to keep one's own accounts straight bears directly on the rise and fall of nations. Modern financial auditing and independent inspection take the same stance. In science, it is preregistration (Nosek, 2018)33 and reproducibility (echoing the replication crisis of Chapter 3). You register your hypothesis and method before you see the data, so the target cannot be moved after the fact.

The unifying idea is to trade "stop the bad thing before the fact" (which takes verification) for "be sure to catch the bad thing after the fact" (which takes only a faithful ledger). The benefit is twofold. Errors can be corrected, and because there is "no getting away with it," the ledger also deters.

This move, too, has a single standard failure mode, and it is extremely common: detection that no one acts on. An audit log no one reads, or an alert everyone ignores, is as good as none. The Equifax data breach of 2017 is a textbook case. A known vulnerability went unpatched for too long. The intruders stayed in the system for roughly seventy-six days before anyone noticed, and the personal information of about 147 million people leaked out. The traces were all there in the logs. No one looked. Detection without response is a sham. (Another hidden danger is that the log itself can be tampered with, which is what the cryptographic methods above are meant to prevent.)

All Eight Moves Together: The End of Part III

Look at this last pair together. Containment shrinks the cost of failure before the fact, and the audit trail makes sure failure is found after it. Neither one tries any longer to make the unverified thing correct. Instead, both change the shape of failure itself: one shrinks the blast radius, and the other moves the check to after the fact.

With that, all eight moves are in place, in four pairs:

  • Compressing the Unknown (Chapter 9): certificate / bound, optimal screening.
  • Borrowed Judgment (Chapter 10): oracle in the loop, redundancy / consensus.
  • Swap the Problem (Chapter 11): proxy substitution, calibration.
  • Contain the Consequences (Chapter 12): containment / fencing, audit trail.

This is the comparison table the book has been building toward, and its main contribution. Across the four settings of Part II and in science, the same moves keep turning up in different jargon, and they are always these eight. As the iron rule of Chapter 4 demands, each move has come, as far as possible, with an account of its mechanism, its forms across fields, and its standard failure mode, rather than resting on surface resemblance alone.

But one sharp question is still open. Why these eight, of all things? Is this a list I have cobbled together, or does each move answer to something more basic that cannot be avoided? If it is only a list, the book is at best a useful classification manual. If there really is a structure behind it, then the "convergence" has finally been explained. Part IV pursues that question. It first tries to hang the eight moves on a common skeleton, then asks directly whether this is a law or a very strong empirical pattern.

References

Waypoints: 1. historical scientific judgment; 2. theoretically studied material; 3. how science progresses; 4. how to live in an unverifiable world. This section was checked source by source.

  1. J. Saltzer & M. Schroeder (1975). The Protection of Information in Computer Systems. Proceedings of the IEEE. doi:10.1109/proc.1975.9939 [2] This survey systematizes the design principles of protection mechanisms, and among them the principle of least privilege becomes the wellspring of the "containment" move: grant a component only the capability it needs to do its own job, and nothing more. The whole line of thought behind this chapter's "shrink the blast radius" originates here, and it is the first required reading for understanding why the scope of failure should be fenced off before the fact.
  2. B. Lampson (1973). "A Note on the Confinement Problem." Communications of the ACM. doi:10.1145/362375.362389 [2] Lampson poses the "confinement problem": how to ensure that a called program cannot leak or misuse the information it has access to, including hard-to-block side channels such as covert channels. It sets a precise problem statement for "caging the thing that goes wrong," which is exactly the core that this chapter's containment move means to address.
  3. D. Bell & L. LaPadula (1973). Secure Computer Systems: Mathematical Foundations. The MITRE Corporation. link [2] The Bell-LaPadula model characterizes confidentiality in a formal way: information may flow only from lower to equal or higher classification, from which comes the famous "no read up, no write down" rule. It demonstrates how the "scope of failure" can be written as a provable lattice structure, a founding work in the theorization of security models.
  4. K. Biba (1977). Integrity Considerations for Secure Computer Systems. The MITRE Corporation. Google Books [2] The Biba model is the dual of Bell-LaPadula, concerned with integrity rather than confidentiality: information may flow only from high trust to low trust, to keep low-trust data from contaminating critical components. Taken together, the two show that the same lattice-theoretic framework can bound the spread of failure from two directions.
  5. D. Denning (1976). "A Lattice Model of Secure Information Flow." Communications of the ACM. doi:10.1145/360051.360056 [2] Denning unifies information-flow security into a single lattice model: data is tagged with security labels, and flow is required to proceed only along the lattice's partial order, thereby constraining where information may go statically, at compile time or run time. It supplies a common mathematical language for the preceding security models, the theoretical core of information flow control.
  6. D. Clark & D. Wilson (1987). "A Comparison of Commercial and Military Computer Security Policies." IEEE Symposium on Security and Privacy. doi:10.1109/sp.1987.10001 [2] Clark and Wilson point out that commercial settings care more about integrity than about military-style confidentiality, and propose an integrity model centered on well-formed transactions and separation of duties. It extends "containment" from military classification levels to everyday settings such as commercial accounting, showing that the form of bounding the scope of failure varies with the domain.
  7. J. Dennis & E. Van Horn (1966). "Programming Semantics for Multiprogrammed Computations." Communications of the ACM. doi:10.1145/365230.365252 [2] This early paper introduces the concept of the capability: access rights attach directly to a reference in the form of an unforgeable token, and only holding the token lets one operate on the object. It is the wellspring of the capability security model, providing a mechanism-level implementation path for "granting only the minimum capability needed."
  8. N. Provos, M. Friedl & P. Honeyman (2003). "Preventing Privilege Escalation." 12th USENIX Security Symposium. link [2] The authors discuss how to use privilege separation to isolate privileged operations into a minimal, trusted slice of code, so that the main program, even if breached, can act only at low privilege; OpenSSH's privilege separation is the representative practice. It brings least privilege down to the engineering detail of real systems.
  9. I. Goldberg, D. Wagner, R. Thomas & E. Brewer (1996). "A Secure Environment for Untrusted Helper Applications." 6th USENIX Security Symposium. link [2] This paper introduces Janus, which uses system-call interception to build a restricted runtime environment for untrusted programs, an early exemplar of the user-space sandbox. This chapter lists the sandbox as an incarnation of the containment move, and this paper is exactly the representative source of the sandbox idea.
  10. C. Perrow (1984). Normal Accidents: Living with High-Risk Technologies. Basic Books. Google Books [2][1] Perrow advances the "normal accidents" thesis: in highly complex, tightly coupled systems, catastrophic accidents are not accidental but structurally inevitable, and cannot be rooted out by adding protections. It supports this chapter's stance from the other side: when errors cannot be prevented one by one, the center of gravity should move to managing consequences rather than vainly trying to abolish failure.
  11. N. Leveson (2011). Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press. doi:10.7551/mitpress/8179.001.0001 [2] Leveson reconstructs safety engineering with systems theory, proposing the STAMP model, which treats an accident as a failure of the control structure rather than the fault of a single part, and stresses using constraints to bound hazardous states. It supplies a systems-level methodology for "fencing off the scope of failure before the fact."
  12. J. Reason (1990). Human Error. Cambridge University Press. doi:10.1017/cbo9781139062367 [2] Reason analyzes human error systematically and proposes the famous "Swiss cheese model": every layer of protection has holes, and only when the holes in multiple layers happen to line up does an accident run straight through. It is the very source of this chapter's defense-in-depth intuition, and also a reminder that once the holes in the layers become correlated, the multiple layers degenerate into one.
  13. E. Hollnagel, D. Woods & N. Leveson (2006). Resilience Engineering: Concepts and Precepts. Ashgate. Google Books [2][4] This collection lays the foundations of "resilience engineering": a system's safety lies not in eliminating faults but in having the capacity to absorb disturbances and to keep running and recover after failure. It accords closely with this chapter's theme, shifting the center of gravity explicitly from "not failing" to "surviving even when it fails."
  14. A. Avizienis, J.-C. Laprie, B. Randell & C. Landwehr (2004). "Basic Concepts and Taxonomy of Dependable and Secure Computing." IEEE Transactions on Dependable and Secure Computing. doi:10.1109/tdsc.2004.2 [2] This widely cited taxonomy clarifies the chain of fault, error, and failure, and the relations among means such as fault tolerance, fault prevention, and fault detection. It provides an agreed terminological framework for the various containment and audit-trail methods this chapter discusses, suitable as a reference for conceptual calibration.
  15. M. Nygard (2007). Release It! Design and Deploy Production-Ready Software. Pragmatic Bookshelf. Google Books [2][4] Nygard writes reliability engineering as a practical handbook, proposing stability patterns such as the circuit breaker, the bulkhead, timeouts, and compartment isolation, to keep a local fault from cascading into a global collapse. The main text takes it as the representative of blast-radius design, a direct read on applying the containment idea to production systems.
  16. R. Anderson (2020). Security Engineering: A Guide to Building Dependable Distributed Systems (third edition). Wiley. doi:10.1002/9781119644682 [2][4] Anderson's monumental work ranges across cryptography, access control, economic incentives, and real-world offense and defense, an authoritative comprehensive textbook of security engineering. Almost every topic this chapter touches, least privilege, auditing, tamper-evidence, can be found there in fuller development, suitable as a base text to read through.
  17. N. N. Taleb (2012). Antifragile: Things That Gain from Disorder. Random House. Google Books [4] Taleb introduces "antifragility": some systems not only survive volatility but benefit from it, the key being to cap the downside and preserve the upside. This chapter cites it to show that the extreme posture of containment is to put a ceiling on loss, a perspective for understanding "contain the consequences" from the angle of risk management.
  18. R. Merkle (1980). "Protocols for Public Key Cryptosystems." IEEE Symposium on Security and Privacy. doi:10.1109/sp.1980.10006 [2] Merkle here proposes the idea of tree-structured authentication using a hash tree (the Merkle tree): a large body of data is merged into a single root hash, and the authenticity of any one item can be checked with only an $O(\log n)$ path. It is the technical bedrock of this chapter's "audit trail" move, and the common ancestor of the later Certificate Transparency and blockchain.
  19. S. Haber & W. S. Stornetta (1991). "How to Time-Stamp a Digital Document." Journal of Cryptology. doi:10.1007/bf00196791 [2] The two authors propose chaining document timestamps together with hashes, so that any later tampering breaks the continuity of the chain and is exposed. This is the pioneering work on chained tamper-evident records, directly inspiring the later blockchain structure, and the key to understanding why the audit trail "cannot be altered."
  20. B. Schneier & J. Kelsey (1998). "Cryptographic Support for Secure Logs on Untrusted Machines." 7th USENIX Security Symposium. link [2] This paper designs a scheme for protecting logs on machines that may be compromised: even if an attacker later gains control, they cannot delete or alter the earlier records without being noticed. It advances tamper-evident logging into untrusted environments, one of the core techniques of this chapter's audit-trail move.
  21. B. Schneier & J. Kelsey (1999). "Secure Audit Logs to Support Computer Forensics." ACM Transactions on Information and System Security. doi:10.1145/317087.317089 [2] This is the journal-version extension of the previous work, treating more fully the construction of secure audit logs that support forensics. It shows that the audit trail must not only record faithfully but also withstand adversarial checking after the fact, matching this chapter's goal of "making the error show itself after the fact."
  22. M. Bellare & S. Miner (1999). "A Forward-Secure Digital Signature Scheme." CRYPTO '99. doi:10.1007/3-540-48405-1_28 [2] Bellare and Miner propose the forward-secure signature: the key evolves periodically, so that even if the current key is leaked, an attacker cannot forge signatures from earlier periods. It provides a key safeguard for the audit trail, so that past records remain unimpersonable and untamperable even after the private key is compromised.
  23. S. Crosby & D. Wallach (2009). "Efficient Data Structures for Tamper-Evident Logging." 18th USENIX Security Symposium. link [2] The authors design a tamper-evident log structure that can be appended to and audited efficiently, letting a verifier confirm the integrity and consistency of records without trusting the log server. It integrates the foregoing cryptographic methods into a deployable data structure, a representative work in the engineering of the audit-trail move.
  24. B. Laurie, A. Langley & E. Kasper (2013). RFC 6962: Certificate Transparency. IETF. doi:10.17487/RFC6962 [2] Certificate Transparency records all issued TLS certificates in a public, append-only Merkle log auditable by anyone, leaving mis-issued or malicious certificates nowhere to hide. It is the real-world model for this chapter's "global audit ledger that anyone can verify," showing how the audit trail can be deployed at scale.
  25. L. Lamport, R. Shostak & M. Pease (1982). "The Byzantine Generals Problem." ACM Transactions on Programming Languages and Systems. doi:10.1145/357172.357176 [2] This classic paper formalizes the Byzantine fault tolerance problem: when some nodes may behave arbitrarily badly, how can the honest nodes agree on a value, and it gives the theoretical bound on how many faulty nodes can be tolerated. It is the consensus-theoretic foundation on which the publicly verifiable ledger rests, listed by this chapter under "theoretically studied material."
  26. M. Castro & B. Liskov (1999). "Practical Byzantine Fault Tolerance." 3rd USENIX Symposium on Operating Systems Design and Implementation (OSDI). link [2] Castro and Liskov give the first Byzantine fault tolerance algorithm, PBFT, practical in a real asynchronous network, bringing theoretical consensus to engineerable performance. It shows that a ledger anyone can verify and that tolerates malicious nodes is no fantasy, providing implementation support for the trusted basis of the audit trail.
  27. S. Nakamoto (2008). Bitcoin: A Peer-to-Peer Electronic Cash System. White paper. link [2] Nakamoto's white paper proposes Bitcoin: a decentralized, append-only blockchain driven by proof of work, letting mutually distrusting parties agree on the transaction history. This chapter views it in essence as a globally verifiable audit ledger, the extreme realization of the audit-trail idea on an open network.
  28. D. Weitzner, H. Abelson, T. Berners-Lee, J. Feigenbaum, J. Hendler & G. Sussman (2008). "Information Accountability." Communications of the ACM. doi:10.1145/1349026.1349043 [2][4] The authors argue for shifting from "blocking access before the fact" to "accountability after the fact": allow information to flow, but require that its uses be auditable and that violations be traceable and answerable. This is wholly isomorphic to this chapter's "audit-trail" posture of moving the check from before to after the fact, a programmatic statement of the idea in privacy governance.
  29. J. Soll (2014). The Reckoning: Financial Accountability and the Rise and Fall of Nations. Basic Books. Google Books [1] Soll argues from financial history that whether a regime can keep and faithfully present its own accounts bears directly on its rise and fall, with double-entry bookkeeping the key accountability technique among them. It traces the lineage of the audit trail back to humanity's earliest tamper-evident ledgers, showing that the power of the audit ledger is of long standing.
  30. B. Beyer, C. Jones, J. Petoff & N. Murphy (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly. Google Books [4] This book introduces Google's SRE practice systematically, including the error budget, canary releases, monitoring and alerting, and controlled failure drills. This chapter borrows its error budget and canary releases to show how containment can be institutionalized in large-scale production, a modern model for the engineering of "contain the consequences."
  31. J. Ioannidis (2005). "Why Most Published Research Findings Are False." PLoS Medicine. doi:10.1371/journal.pmed.0020124 [3] Ioannidis argues with statistical modeling that under conditions of low priors, small samples, multiple comparisons, and excessive researcher degrees of freedom, a great many published findings are likely false positives. This is an analytical argument rather than an empirical replication study, supplying a problem diagnosis for the preregistration and reproducibility this chapter mentions.
  32. Open Science Collaboration (2015). "Estimating the Reproducibility of Psychological Science." Science. doi:10.1126/science.aac4716 [3] This is a large-scale empirical effort: many groups of researchers attempt to replicate over a hundred psychology studies, and a sizable share fail to reproduce. It turns Ioannidis's theoretical worry into visible data, the landmark evidence for the "replication crisis," echoing this chapter's emphasis on after-the-fact verifiability.
  33. B. Nosek, C. Ebersole, A. DeHaven & D. Mellor (2018). "The Preregistration Revolution." PNAS. doi:10.1073/pnas.1708274114 [3] Nosek and colleagues advocate preregistration: publicly registering hypothesis and analysis method before seeing the data, separating exploratory from confirmatory research so the target cannot be moved after the fact. It is the "audit trail" practice in science, shifting verification from trust before the fact to checking after it, corresponding directly to this chapter's theme.

Open with WeChat

Scan this page in WeChat, then use WeChat's share menu.