Recurrence and Correction

Why failures return, and what it takes to stop them

The Problem of Recurrence

The same failure returns. An institution investigates a disaster, publishes the findings, implements the reforms, and years later produces a structurally identical disaster. An industry pays for a lesson, learns it, and relearns it a generation on. A regulatory framework is built to prevent a crisis, erodes across decades, and the crisis returns. The pattern is so common that it is often treated as inevitable, a fact of institutional life rather than a thing that calls for explanation.

It calls for explanation. When the same structural failure returns after a system has already studied it, something specific has happened. The recurrence is not bad luck and it is not mysterious. It is the trace of where the previous correction landed.

A failure that returns is telling you something about the repair that was supposed to prevent it.

Recurrence Is Evidence

This is the first move, and everything else on the page is built on it. A returning failure is information, and the information has an admission condition: the return must be shown to share the relevant generating condition with the original. A similar-looking outcome can arrive through a different mechanism, a changed environment, or an independent failure, and reports nothing about the earlier repair. Structural recurrence is the case where reconstruction establishes the same generating condition participating again. Where that is established, the reading follows: the prior correction did not reach that condition, or reached it and did not persist. The correction may have been thorough. It may have been sincere. It may have changed procedures, restructured organizations, and retrained personnel. The verified return reports that whatever it changed was not the condition, or did not stay changed.

This reframes what recurrence is. The documented cases show the lesson was usually studied with great care. The verified recurrence is evidence about depth: it marks the distance between where the correction acted and where the failure lived.

Not every recurrence follows the same path. Some corrections never reach the condition that generated the failure. Others reach it, hold for a time, and later degrade. The first is a failure of arrival. The second is a failure of persistence. Both produce recurrence, but they report different things about the correction that preceded them.

Read the returnA verified recurrence is a measurement. Verification means the reconstruction shows the same generating condition participating in the return, not merely a similar outcome. So established, it reports that the prior correction failed to reach, or failed to persist at, the level where the failure was generated. Sometimes the correction acted too shallowly. Sometimes it reached the right level and later decayed.

The Levels of Correction

Corrections differ by depth, not by effort or intelligence. A correction can act on a symptom, a procedure, a policy, a structure, or the invariant condition beneath all of these, the condition whose absence is what generated the failure in the first place.

At the symptom level, a correction stops the specific event. At the procedure or policy level, it constrains a class of similar events, for as long as the procedure is followed. At the structural level, it changes what the system can do. At the invariant-condition level, it restores the thing whose absence made the failure possible. A correction stops the pattern when it reaches the level where the generating condition actually resides: sometimes a design, sometimes a structure, sometimes a policy, sometimes the invariant condition underneath them all. The depth required is set by where the failure was generated, and a correction above that level reaches the instance.

The reason recurrence is so common is that the deeper levels are harder to reach, more expensive to restore, and easier to mistake for already-done. A reform that updates procedures and reorganizes reporting lines looks like structural correction. It produces documents, announcements, and a record of action. It can be complete and sincere and still act entirely above the level where the failure was generated. The system then shows the appearance of correction and the unchanged condition at once.

Some corrections stop an event. Some corrections stop a pattern. The difference is whether the correction reaches the level where the failure was generated.

What the Cases Show

The Case Verification corpus was built one failure at a time, each paper reconstructing a single documented event. Read across, the papers describe recurrence with unusual precision, because each one identifies the level at which the prior correction acted and the level at which the failure lived. The reconstruction is what qualifies each return as structural recurrence in this page's sense: the same generating condition, shown participating again.

Challenger and Columbia. The same institution produced two structurally identical losses seventeen years apart. The investigation after the first correctly identified the structural conditions. The reforms acted on organization and process. The second loss reproduced the first inside the same architecture, with the conditions the first investigation had named still present. Recurrence inside a corrected system. (CV-002 Challenger, CV-011 Columbia)

Texas City and Macondo. A comprehensive independent assessment identified the operator's structural failure conditions and documented their presence across multiple facilities. The corrective actions were procedural and organizational. Five years later the same operator produced a catastrophe exhibiting the same underlying structural conditions in a different domain. Recurrence after the assessment that named the conditions. (CV-006 BP Texas City, CV-008 Deepwater Horizon)

Banking regulation. A constraint architecture was installed after a systemic crisis, eroded across decades through statutory carve-out and institutional memory decay, and the crisis returned roughly seventy-five years on. The recurrence came from the slow loss of the conditions the original constraint had supplied, no single decision behind it. Recurrence through the decay of memory. (CV-013 Banking Regulation Cycles)

Supply-chain compromise. A trust architecture accepted signed software as authenticated without verifying that the signed build had not been altered. The technique returned across separate events because the structural condition, the unverified trust relationship, was never restored. Recurrence through an unchanged structural condition. (CV-005 SolarWinds SUNBURST)

In each case the correction was real and the failure returned, because the correction either acted above the level that generated it or failed to persist there. These are not stories about negligence. They are measurements of depth and persistence.

Why Some Systems Learn

Most of the corpus asks why correction failed. One case asks the opposite question, and it is the more revealing one. Apollo 13 is a system that suffered an acute, multi-system failure and did not produce a catastrophe. The corrective authority formed in real time, the system's reading of its own state was accurate as the event unfolded, and the failure was arrested before nonlinear transition. There is no wreckage to walk.

This is what makes the Apollo 13 and Columbia pair the sharpest object in the corpus. The two events run on substantially the same institutional architecture. They reach opposite outcomes. The architecture was the same. The conditions inside it were what differed: whether the authority that could correct the failure was able to form around the signal in time, whether the signal reached the level that could act on it, and whether the structure was free of the pressures that push authority away from correction. In Apollo 13 those conditions were present. In Columbia, seventeen years of accumulated load had degraded them. That split, between a system that could read and move in time and one that read but could not move on the reading, is decomposed formally across all thirteen cases in ER-002 — Three-State Decomposition of the SAG.

The corpus records one further observation about Apollo 13, and it is a different finding from the arrest itself. The arrest was a real-time event: corrective authority forming during the emergency. What came after was separate: the correction engaged the failure mechanism directly, at the design level. Recurrence of that failure does not appear across the four subsequent missions, a small sample the corpus is careful to note, and one bounded by how few missions the program had left. It is the counterexample that makes the question concrete: most cases show correction acting above the failure and the failure returning; this one shows correction reaching the mechanism and the failure not returning. (CV-010 Apollo 13)

The inverted questionWhy did this correction succeed? Because the authority to correct was able to form, the signal reached the level that could act, and the repair engaged the mechanism itself. Stated as one instance, left open to what further cases show.

Reality's Role

Put the two halves together and the logic arrives somewhere specific. Recurrence is what becomes available when a system fails to establish or fails to keep the capacity to correct departures from reality. Correction is what becomes possible when that capacity is preserved or rebuilt. The returning failure and the arrested one are the same mechanism seen from opposite sides.

This is the second half of a claim the corpus has been making elsewhere. Reality is the standing measure a system drifts from, and what makes the drift correctable at all. The capacity to correct is the system's to preserve or lose. The Case Verification record adds the consequence: a system that fails to preserve that capacity retains the conditions that generate its failures, and recurrence is one consequence by which the loss becomes visible under renewed load. The repetition, when it arrives, is the shadow that a lost correction casts.

Reality is the standing measure a system drifts from. A system that loses the capacity to correct retains the conditions that generate its failures.

Recurrence and correction are not two topics. They are one structure read in two directions. Recurrence is the phenomenon that demands the explanation. Correction is the object the explanation is about. Everything else the corpus studies, signal normalization, the conditions for corrective authority, the depth at which a repair acts, the decay of institutional memory, is a way of describing where, between the failure and its return, the capacity to correct was kept or lost.

The Case Verification corpusEach case named above is reconstructed in full in its own paper. The corpus was built one failure at a time, and the cross-case synthesis reads across all thirteen for the structural regularities they share. The complete set is below, for the reader who wants to follow recurrence and correction down to the documented record.

Cases discussed on this page:
CV-002 Challenger · CV-011 Columbia · CV-006 BP Texas City · CV-008 Deepwater Horizon (Macondo) · CV-013 Banking Regulation Cycles · CV-005 SolarWinds SUNBURST · CV-010 Apollo 13

The rest of the corpus:
CV-001 Air France 447 · CV-003 Therac-25 · CV-004 Knight Capital Group · CV-007 Global Financial Crisis · CV-009 Boeing 737 MAX · CV-012 Vasa

Neighboring Work

The questions next door.

Chris Argyris & Donald Schön · The two loops

The depth distinction has a canonical prior statement in organizational learning. Chris Argyris and Donald Schön separated two ways an organization corrects a detected error. In single-loop learning, the correction changes the action and leaves the governing variables, the goals, values, and rules behind the action, exactly as they were; their image for it is a thermostat, competent forever at one question. In double-loop learning, the error reaches the governing variables themselves: they are examined, altered, and only then are the actions changed. Much of their corpus examines why organizations reach the first loop so much more easily than the second.

The kinship with this page's levels is direct, and the grant is easy: the claim that corrections differ by the depth at which they act has its nearest ancestor in that distinction. The two frameworks divide at what the depth is made of. For Argyris and Schön, the deeper loop reaches the governing variables, and those can live anywhere the organization's theory-in-use lives: in values and norms, and equally in policies, structures, and routines. The boundary is the test. Double-loop learning asks whether the governing variables behind action were revised. The reading here asks whether the correction reached, and persisted at, the condition that generated the return. A verified return under renewed load can show that it did not; a nonreturn, on its own, shows less. Recurrence, on this page, is what measures the depth after the fact.

James Reason · The latent condition

Safety science located where failure lives. James Reason distinguished active failures, the unsafe acts at the sharp end whose effects are direct and short-lived, from latent conditions: decisions and designs that enter a system high up and far upstream, then reside in it, sometimes for years, until they combine with local circumstances to breach every defense at once. He called them resident pathogens, and the model built on them, layered defenses whose holes occasionally align, became one of the most widely used pictures in accident analysis.

The page's central observation, that a correction can be real and still act above the level where the failure was generated, has its safety-science ancestor here: a repair that addresses the act while the condition goes on residing is Reason's anatomy restated as a correction problem. His corpus centers on how organizational accidents are produced, the trajectory from upstream decision through defenses to event. The object here begins after that story has been told once: what the failure's return, following a studied and sincere correction, reports about the repair. The cases above are read that way, each return a report on the level a correction reached, or failed to keep.

Charles Perrow · The normal accident

The strongest case for treating recurrence as inevitable belongs to Charles Perrow. Normal Accidents argued that in systems combining interactive complexity with tight coupling, small failures will sometimes interact in ways no designer anticipated and no operator can diagnose in time, and that accidents of this kind are a property of the system itself: normal, in his word, not because they are frequent but because the system's own characteristics produce them. The argument is structural to its core. It names two inspectable system properties, and it warns that a favorite remedy, redundancy, can add the very complexity it is meant to defend against.

This page opens against the position Perrow gave its strongest form, so the boundary matters. His thesis conditions accident production on system structure: what a system is like determines what it will sometimes do. The claim here conditions recurrence on correction: where a studied failure returns, the return reports on the depth and persistence of the repair. The two claims meet at the cases. Whether restoring an invariant condition changes what returns, even in territory with Perrow's two properties, is exactly what the Case Verification corpus exists to test, one documented failure at a time, and the corpus treats that answer as open where the record has not yet closed it.

Argyris & Schön, Organizational Learning: A Theory of Action Perspective (Addison-Wesley, 1978) · Reason, Human Error (Cambridge University Press, 1990); Managing the Risks of Organizational Accidents (Ashgate, 1997); "Human error: models and management," BMJ 320 (2000), 768–770 · Perrow, Normal Accidents: Living with High-Risk Technologies (Basic Books, 1984; updated edn, Princeton University Press, 1999)

Related materials: Case Verification · CV-SYN-001 Cross-Case Structural Analysis · ER-002 Three-State Decomposition of the SAG · Reality Contact · Restorative Realism