Chapter 8: The Organization Blind to Itself
Thesis: A large organization or a state cannot directly observe its own knowledge and activity, which are distributed, partly hidden, and at times strategically concealed, so it reaches for a proxy it can see: a metric. This is where proxy substitution's Goodhart failure is most glaring, and the organization then shores the proxy up with auditing (the audit trail) and redundancy.
A Giant That Cannot See Itself
The principal-agent problem from the chapter on the released agent now grows in scale. The principal is no longer a single person but an entire organization, a whole state. The people entrusted to act are thousands upon thousands, scattered everywhere. The result is a new and almost absurd situation: this giant cannot see itself clearly.
It wants to know how many people it has, what is being planted, who is doing what, and how well. Yet none of this knowledge sits anywhere it can read directly. This chapter asks what an organization does when the thing to be verified is its own knowledge, which is distributed, hidden, and apt to dodge being seen. Here several of the earlier predicaments combine. There is partial observability (knowledge scattered at the edges), plus the adversarial face (the people being watched will turn around and manipulate what is being watched).
Distributed Knowledge
Hayek's 1945 essay "The Use of Knowledge in Society"1 set out the problem. The knowledge a society runs on is never concentrated in any one place. It is dispersed among countless individuals. It is local knowledge, about a particular time and a particular place, and often it cannot even be put into words. Which machine has a small fault today, which customer is quietly about to drift away, which side path will collapse after the rain: the people who hold such knowledge often do not realize it counts as "knowledge," let alone find a way to package it up and hand it to the center. Polanyi called this layer the tacit dimension2: we know far more than we can tell.
So the organization faces more than information that "has not been collected yet." Even if everyone were loyal and cooperative, that local, tacit knowledge would still evaporate in the course of being gathered. The "whole picture of the organization" that the center wants cannot, in principle, be fitted faithfully into any container that could verify it. This is partial observability at the scale of a society, and it comes with a harder floor: such knowledge is local by its very nature, and it cannot be gathered into any one place.
The Urge Toward Legibility
If you cannot see something clearly, the impulse is to make it easier to see. Scott's 1998 book Seeing Like a State3 gives this urge a name: legibility. Before a state can act on society, it must first remake society into a shape it can read. It surveys the land and draws cadastral maps. It imposes fixed surnames on people who once had only nicknames, bynames, or patronymics. It standardizes weights and measures, and it rolls out standardized scientific forestry. These are not neutral acts of recording. They reshape reality itself so that reality will fit the table. Taken together, Hacking's The Taming of Chance32, Desrosières's The Politics of Large Numbers31, and Bowker and Star's Sorting Things Out30 form a history of "making society countable."
The danger of legibility is that the map has to simplify, and once an organization acts only by the map, whatever the map has erased comes back to bite. Scott's most forceful case is scientific forestry. To make the forest "legible, countable, and taxable," the Prussians turned tangled natural woodland into uniform single-species plantations that were easy to inventory. The first generation grew splendidly. By the second, the soil was exhausted, pests had spread, and the forest was dying off in swaths, so badly that German even coined a word for it: Waldsterben, the death of the forest. The cleaner the map, the more deadly the loss of the local knowledge it erased, the knowledge that had kept the system running. Here the organization manufactures the observability it lacks, and the price is that it flattens, with its own hands, the very complexity that let it run.
The Proxy Metric and Its Goodhart Collapse
The urge toward legibility most often ends in a metric. The things an organization really cares about (health, learning, productivity, public welfare) cannot be observed directly. So it grabs the proxy it can see: the KPI, GDP, exam scores, citation counts, emergency-room waiting times.
This is the proxy substitution we met in Chapter 7. Here, though, it fails in the opposite way, and that contrast is one of this book's main threads. The mathematician's proxy is faithful but no easier: an equivalent reformulation really is equivalent, but it has not become any easier to solve. The organization's proxy is the reverse, easier but unfaithful. The metric is easy to measure, of course, but its link to the true target snaps the moment the metric itself becomes the target.
This break goes by many names. Goodhart, in 19754: once a metric is made a policy target, its reliability as a metric falls apart. Campbell, in 19796, stated the social version of the same thing. As early as 1956, Ridgway had catalogued the "dysfunctional consequences of performance measurements"7, and Kerr's 1975 essay "On the Folly of Rewarding A, While Hoping for B"8 made the idea part of management common sense. Strathern put it in its most distilled form5: when a measure becomes a target, it ceases to be a good measure.
A deeper layer is reactivity. In 2007 Espeland and Sauder12 pointed out that a public ranking is remaking the world, not describing it. A ranked university will change itself to fit the ranking's formula, so what the metric "measures" is the very behavior it has called into being. Bevan and Hood11 documented the gaming of metrics inside the English health system. In 1995 Smith10 analyzed how publishing performance data invites a string of unforeseen consequences. And Merton's 1936 paper on "the unanticipated consequences of purposive social action"9 is the source from which all of this flows.
Such collapses are common, and the cost is sometimes staggering. To hit a "cross-selling" metric for accounts, Wells Fargo employees secretly opened large numbers of fake accounts without customers' knowledge. When the affair came to light in 2016, the estimate was some 2 million accounts (later revised upward to about 3.5 million). The bank was fined 185 million dollars, and more than five thousand employees were fired. The number the bank had enshrined had destroyed the very customer relationships it was meant to measure. An earlier case, almost a parable, comes from colonial-era Delhi. To get rid of snakes, the authorities offered a bounty for dead cobras, so residents started breeding cobras to collect it. When the bounty stopped, the snakes were all set free, and the snake problem was worse than before. That is where the "cobra effect" gets its name.
Why is a proxy bound to be distorted? Principal-agent theory gives the rigorous explanation. Holmström's 1979 informativeness principle14 says that rewards should be tied to signals that carry information about "effort." But once effort has many dimensions and you can measure only a few of them, trouble begins. Holmström and Milgrom's 1991 multitask analysis15 (the multitask principal-agent model) spells it out. When a person has to attend to measurable and unmeasurable tasks at once, the more heavily you reward the measurable part, the more they will shift effort away from the unmeasurable part and toward the measurable one. Let the true target be $G$ and the observable proxy be $P$. Under the status quo the two are correlated. The problem is that this correlation is a product of behavior, not an objective law. Once $P$ is made the target of pressure,
$$\arg\max_{a} P(a)\ \quad\text{vs.}\quad\ \arg\max_{a} G(a),$$
the rational agent goes looking for actions that raise $P$ while doing nothing for $G$, or even harming it. The pressure to optimize crushes the correlation. The teacher who teaches to the test, the hospital that schedules patients to lower one particular waiting-time figure, the researcher who slices one paper into the smallest countable units of publication: all of them are the same mechanism at work.
Shoring It Up With Auditing and Redundancy
On its own, the proxy will collapse, so the organization adds two more moves. These, too, recur throughout this book.
The audit trail and auditing. Double-entry bookkeeping is one of humanity's oldest audit chains. In The Reckoning22, Soll argues that the ability to keep accounts that can be checked bears directly on the rise and fall of one nation after another: those that can reckon with themselves are the ones that endure. Modern financial auditing and independent inspection all come down to trading "fraud cannot be prevented in advance" for "fraud can be detected after the fact." But this move has an ailment of its own. Power's 1997 book The Audit Society20 spells it out. When verification itself becomes a ritual, what the organization produces is the appearance that "everything is under control," not control itself. Shore and Wright's "audit culture"17 and O'Neill's reflections on "trust" in the 2002 Reith Lectures21 describe the same alienation. To be held accountable, institutions pour enormous effort into manufacturing traces that can be inspected, while the real work gets pushed to one side.
Redundancy and consensus. Landau's underrated 1969 article16 rehabilitated "duplication and overlap." In a system where no part is fully reliable, redundancy is not waste but a source of reliability: several independent checks are harder to fool all at once than a single authority. The move holds only under one condition, independence, which the next part will stress again and again. If several checks actually share a common origin, a correlated failure will break the central promise of redundancy at a single stroke.
What Comes Next: The Close of Part II
That completes our tour of the four settings. The person at the console, the agent set loose, the mathematician at the wall, and the organization blind to itself face very different sources of unverifiability: preferences hidden in someone's heart, future behavior in an open world, statements undecidable in principle, and knowledge that is distributed and apt to dodge being seen. Yet what they reach for is the same small set of things.
Most worth setting side by side are the two opposite failures of proxy substitution. The mathematician comes to grief on "faithful but no easier." The organization comes to grief on "easier but unfaithful." Both ends of that 2×2 table in Chapter 7 now have flesh on them. They are not two moves but two directions in which a single move can fail. A good proxy has to avoid both ends at once, being faithful and easier at the same time, and that is so rare that it is nearly the whole of the craft. Chapter 11 will formally join these two ends. The principal-agent skeleton, too, has grown from a snippet of code in Chapter 6 into a whole state here.
Part II has now shown the moves as they appear in practice: embedded in their settings and tangled together. They are scattered, they change names, and they are mixed into each field's own jargon. Part III sets out to pull each move out of the field where it grew, strip it down, give it a name of its own, and cover every setting at once. That comparison table is the heart of this book.
References
Waypoints: 1. historical scientific judgment; 2. theoretically studied material; 3. how science progresses; 4. how to live in an unverifiable world. This section was checked source by source.
- F. A. Hayek (1945). "The Use of Knowledge in Society." American Economic Review, 35(4), 519-530. link [2][4] Hayek argues that the knowledge on which a society runs is never concentrated in one place but dispersed among countless individuals, that it is local knowledge concerning a particular time and a particular place, and that it cannot be faithfully gathered by any center. This essay is the direct point of departure for this chapter's section "Distributed Knowledge," and it sets the epistemological coloring of the predicament that "the organization cannot see itself clearly."
- M. Polanyi (1966). The Tacit Dimension. Doubleday. Google Books [2] Polanyi proposes the tacit dimension of knowledge, his famous line being "we know more than we can tell." The book uses it to show that a considerable part of the local knowledge dispersed at the edges simply cannot be put into words and handed up, which is a harder floor beneath the organization's difficulty in verifying itself.
- J. C. Scott (1998). Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed. Yale University Press. Google Books [2][4] Scott proposes the concept of "legibility": in order to act upon society, a state uses such means as cadastral maps, fixed surnames, and standardized weights and measures to remake society into a shape it can read, and this simplification often erases the local knowledge that keeps the system running, leading to failures like scientific forestry. This chapter's section "The Urge Toward Legibility" is built precisely on this; it is the core reading for understanding why an organization sets about leveling complexity with its own hands.
- C. A. E. Goodhart (1975). "Problems of Monetary Management: The U.K. Experience." Papers in Monetary Economics, Vol. I. Reserve Bank of Australia. [2] Goodhart was originally speaking of monetary policy, yet he gave an insight later cited everywhere: once a statistical regularity is taken as the target of policy control, its original regularity falls apart. This is the source of the name of this chapter's section "The Proxy Metric and Its Goodhart Collapse," and the starting point for understanding how a proxy is crushed by the pressure to optimize.
- M. Strathern (1997). "Improving Ratings: Audit in the British University System." European Review, 5(3), 305-321. doi:10.1002/(sici)1234-981x(199707)5:3<305::aid-euro184>3.0.co;2-4 [2][4] Strathern, drawing on the experience of auditing in British universities, left Goodhart's law its most distilled popular formulation: when a measure becomes a target, it ceases to be a good measure. This chapter quotes the sentence directly; it is also the single best line for explaining the abstract collapse of the proxy to a reader.
- D. T. Campbell (1979). "Assessing the Impact of Planned Social Change." Evaluation and Program Planning, 2(1), 67-90. doi:10.1016/0149-7189(79)90048-x [2][4] Campbell, from the standpoint of social-science evaluation, proposed "Campbell's law," isomorphic to Goodhart's: the more a quantitative social indicator is used for social decision-making, the more it is subject to corruption pressures, and the more it will distort the very social process it was meant to monitor. This chapter uses it to corroborate that the collapse of the proxy is not peculiar to economics but the same phenomenon discovered again and again across disciplines.
- V. F. Ridgway (1956). "Dysfunctional Consequences of Performance Measurements." Administrative Science Quarterly, 1(2), 240-247. doi:10.2307/2390989 [2][4] Ridgway, very early on, systematically catalogued the dysfunctional consequences of performance measurement, distinguishing the distortions brought by single, composite, and multiple measures. This chapter uses it to show that the gaming and distortion of metrics is a rather old problem, discovered quite early, and not some recent coinage of management studies.
- S. Kerr (1975). "On the Folly of Rewarding A, While Hoping for B." Academy of Management Journal, 18(4), 769-783. doi:10.2307/255378 [2][4] Kerr lists a wealth of real-world examples to show that organizations often reward one kind of behavior while hoping for another that they have not rewarded, with results that naturally run counter to intent. This essay wrote the mismatch of proxy and incentive into the common sense of management, and is the classic source of this chapter's "reward A while hoping for B" mechanism.
- R. K. Merton (1936). "The Unanticipated Consequences of Purposive Social Action." American Sociological Review, 1(6), 894-904. doi:10.2307/2084615 [2][4] Merton systematically analyzed why purposive social action always brings unanticipated consequences, and sorted out their causes, such as ignorance, error, and the imperious immediacy of value. This chapter treats it as the wellspring of a whole series of "unforeseen" phenomena, such as the gaming of metrics and the backlash of legibility.
- P. Smith (1995). "On the Unintended Consequences of Publishing Performance Data in the Public Sector." International Journal of Public Administration, 18(2-3), 277-310. doi:10.1080/01900699508525011 [2][4] Smith classified and sorted out the string of unintended consequences invited by the public release of performance data in the public sector, such as tunnel vision, myopia, misrepresentation, measure fixation, and gaming. This chapter uses it to break the vague "the metric gets distorted" into several recognizable, concrete modes of failure.
- G. Bevan & C. Hood (2006). "What's Measured Is What Matters: Targets and Gaming in the English Public Health Care System." Public Administration, 84(3), 517-538. doi:10.1111/j.1467-9299.2006.00600.x [2][4] Bevan and Hood empirically documented the various ways of gaming metrics in the English National Health Service under "targets and terror" governance, such as scheduling patients so as to lower waiting times, a practice that appeases the metric while doing nothing for real health. This chapter takes it as field evidence of how the gaming of metrics actually happens in public services.
- W. N. Espeland & M. Sauder (2007). "Rankings and Reactivity: How Public Measures Recreate Social Worlds." American Journal of Sociology, 113(1), 1-40. doi:10.1086/517897 [2][4] Espeland and Sauder, using law-school rankings as their example, propose "reactivity": a public measure does not merely describe the world but turns around to reshape the behavior of those being measured, so that what the metric finally measures is the very reaction it has itself called into being. This chapter's paragraph "A deeper layer is reactivity" comes from here; it pushes the failure of the proxy to the level where the metric manufactures reality.
- M. Sauder & W. N. Espeland (2009). "The Discipline of Rankings: Tight Coupling and Organizational Change." American Sociological Review, 74(1), 63-82. doi:10.1177/000312240907400104 [2][4] This companion piece draws on Foucault's concept of discipline to analyze how rankings become embedded in organizations: institutions once loosely coupled are forced under the pressure of rankings into tight coupling, and the external measure is internalized as everyday self-surveillance and organizational change. It complements the previous entry, which treats the mechanism of reactivity, while this one treats how rankings remake an organization's internal structure.
- B. Holmström (1979). "Moral Hazard and Observability." The Bell Journal of Economics, 10(1), 74-91. doi:10.2307/3003320 [2][4] Holmström proposes the informativeness principle: under moral hazard, the optimal reward contract should hang on all signals that carry information about the agent's effort. This chapter uses it to give a rigorous principal-agent explanation of "why the proxy is bound to be distorted," and to lead into the trouble that arises when effort is multidimensional and only a few dimensions can be measured.
- B. Holmström & P. Milgrom (1991). "Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design." The Journal of Law, Economics, and Organization, 7(Special Issue), 24-52. doi:10.1093/jleo/7.special_issue.24 [2] The multitask principal-agent model shows that when a person must attend to measurable and unmeasurable tasks at once, the more heavily the measurable part is rewarded, the more they will draw effort away from the unmeasurable part. This chapter argues from it the mechanism of the proxy's collapse: pressing on the observable metric rationally induces the agent to abandon work that is hard to measure yet truly important.
- M. Landau (1969). "Redundancy, Rationality, and the Problem of Duplication and Overlap." Public Administration Review, 29(4), 346-358. doi:10.2307/973247 [2][4] Landau rehabilitated the "duplication and overlap" so often denounced as waste: in a system whose parts are none of them fully reliable, redundancy is precisely the source of reliability, and several mutually independent checks are harder to fool all at once than a single authority. This chapter's section "Shoring It Up With Auditing and Redundancy" adopts this argument directly, and stresses that its precondition is the mutual independence of the checks.
- C. Shore & S. Wright (1999). "Audit Culture and Anthropology: Neo-Liberalism in British Higher Education." The Journal of the Royal Anthropological Institute, 5(4), 557-575. doi:10.2307/2661148 [2][4] Shore and Wright, taking British higher education as their example, propose "audit culture": under neoliberal governance, the logic of accountability and audit seeps into academic institutions, turning peers into objects of surveillance and reshaping the way people govern themselves. This chapter uses it to show how auditing is alienated from a tool into a culture that makes people spend themselves on manufacturing inspectable traces.
- J. Z. Muller (2018). The Tyranny of Metrics. Princeton University Press. doi:10.23943/9781400889433 [4] Muller, writing for the general reader, surveys the distortions and costs brought by overreliance on quantitative metrics in fields such as medicine, education, policing, and business, and offers judgments about when measurement should and should not be used. The book is a popular synthesis that explains the collapse of the proxy to practitioners, suitable for the reader as an introduction and a point of comparison.
- T. M. Porter (1995). Trust in Numbers: The Pursuit of Objectivity in Science and Public Life. Princeton University Press. doi:10.1515/9781400821617 [2][4] Porter argues that reliance on quantification often springs from a kind of "mechanical objectivity": in situations lacking trust and demanding outward accountability, numbers are used as a tool to suppress personal judgment and ward off challenge. The book provides a deep sociological explanation of why organizations cling to legible numbers, serving as background to both this chapter's sections on legibility and on auditing.
- M. Power (1997). The Audit Society: Rituals of Verification. Oxford University Press. Google Books [2][4] Power points out that when verification itself becomes a set of rituals, what the organization produces is often the appearance that "everything is under control," rather than control itself, and society remakes itself in turn so as to be auditable. This chapter's section "Shoring It Up With Auditing and Redundancy" draws on it to spell out the ailment that the audit move carries within: the more traces, the more the real work gets pushed aside.
- O. O'Neill (2002). A Question of Trust: The BBC Reith Lectures 2002. Cambridge University Press. Google Books [4] O'Neill, in this set of Reith Lectures, reflects on the contemporary culture of accountability: the various measures of transparency and audit meant to rebuild trust often erode the very trust they were intended to foster, leaving people busy coping with inspection rather than doing the work well. This chapter cites it alongside "the audit society" to show how excessive accountability backfires.
- J. Soll (2014). The Reckoning: Financial Accountability and the Rise and Fall of Nations. Basic Books. Google Books [1][4] Soll, with double-entry bookkeeping as his thread, argues that the ability to keep accounts that can be checked bears directly on the rise and fall of one nation after another: those that can reckon themselves are the ones that endure. This chapter uses it to support the claim that "the audit trail and auditing" is one of humanity's oldest chains of verification.
- J. G. March & H. A. Simon (1958). Organizations. John Wiley & Sons. Google Books [2] March and Simon laid the foundations of modern organization theory: the rationality of an organization's members is bounded, and the organization copes with the limits of individual cognitive capacity precisely through division of labor, procedures, and information channels. The book provides a basic framework for "the organization cannot see itself," and is a classic source for understanding how information flows and decays through a hierarchy.
- H. A. Simon (1947). Administrative Behavior: A Study of Decision-Making Processes in Administrative Organization. Macmillan. Google Books [2] Simon proposes bounded rationality, understanding the organization as a set of structures that help members make decisions under limited cognitive capacity. The book is the source for understanding why an organization must rely on simplification, routine, and proxies to run, and it lays the theoretical bedrock for this chapter's account of the limits of organizational self-knowledge.
- R. M. Cyert & J. G. March (1963). A Behavioral Theory of the Firm. Prentice-Hall. Google Books [2] Cyert and March propose a behavioral theory of the firm, stressing that organizational decisions are governed by standard operating procedures, limited search, and the negotiation of goals among parties, rather than by pure optimization. The book helps in understanding the plurality and tension of goals within an organization, and is an important support for this chapter's treatment of the organization as a bounded-rationality actor.
- O. E. Williamson (1975). Markets and Hierarchies: Analysis and Antitrust Implications. Free Press. Google Books [2] Williamson, starting from transaction costs, explains why some activities are coordinated by the market and others are folded into a hierarchical organization: bounded rationality and opportunism make some transactions more efficient to complete within a hierarchy. The book provides an economic explanation of why an organization takes scattered activities under its own roof, and thereby shoulders the difficulty of verifying them.
- K. J. Arrow (1974). The Limits of Organization. W. W. Norton. Google Books [2][4] Arrow concisely explores the organization as a means of coping with the scarcity of information and with uncertainty, and the inherent limits it meets in authority, responsibility, and trust. The book points out that trust is an indispensable lubricant of social functioning that cannot be bought by contract, in distant resonance with the cost of verification examined in this chapter's sections on auditing and redundancy.
- M. Lipsky (1980). Street-Level Bureaucracy: Dilemmas of the Individual in Public Services. Russell Sage Foundation. Google Books [2][4] Lipsky points out that street-level bureaucrats such as teachers, police, and social workers exercise a great deal of discretion under conditions of scarce resources, and that their everyday coping in fact shapes how public policy actually lands. The book is an important reference for understanding why the local knowledge and discretion at the organization's edges is hard for the center to observe and verify.
- J. Q. Wilson (1989). Bureaucracy: What Government Agencies Do and Why They Do It. Basic Books. Google Books [2][4] Wilson examines in detail the actual workings of government agencies, distinguishing types of agency by whether outputs and outcomes are observable, and explains why the real effectiveness of many public agencies is hard to measure. The book provides rich real-world material for this chapter's "the organization blind to itself," and is especially helpful for understanding why proxy metrics are particularly apt to distort in the public sector.
- G. C. Bowker & S. L. Star (1999). Sorting Things Out: Classification and Its Consequences. MIT Press. doi:10.7551/mitpress/6352.001.0001 [2][4] Bowker and Star examine how classification systems silently embed themselves into infrastructure, and how they shape the very reality they meant to record neutrally, with the differences flattened by classification often carrying real consequences. This chapter sets it alongside Hacking and Desrosières, gathering it into the history of "making society countable," to show that classification is an invisible link in the engineering of legibility.
- A. Desrosières (1998). The Politics of Large Numbers: A History of Statistical Reasoning (trans. C. Naish). Harvard University Press. Google Books [2] Desrosières traces the history of statistical reasoning, showing that statistical categories took shape in step with state administration, and that numbers are at once a tool for knowing society and a political act that constructs social reality. This chapter brings it into the genealogy of "making society countable," revealing the provenance of the statistical apparatus behind legibility.
- I. Hacking (1990). The Taming of Chance. Cambridge University Press. doi:10.1017/cbo9780511819766 [2] Hacking examines the rise of statistical and probabilistic thought in the nineteenth century, arguing that the mass collection of population data "tamed chance" and gave birth to concepts such as "the normal" and "normalcy" that govern modern governance. This chapter cites it to show that making society countable is itself a stretch of history that remade cognition, and not a neutral act of recording.