Background·Concepts Versioned reference · undefined format

Contagion

A plain-language tutorial: the two reasons failures arrive together (shared exposure and true propagation), why real disasters are usually both, and why the fix for one makes the other worse.

Concept
background/contagion
Version
v1.0-draft minor1.0 · 2026-07-28 · first published version: four interactive figures built and wired in
minor0.2 · 2026-07-28 · full prose across all seven sections; references added; four figures specified but not yet built
minor0.1 · 2026-07-28 · outline draft: section architecture, core definitions, figure specs
Issued
2026-07-28
Author
Matthew Tanti
Introduction and notation two mechanisms · one confusion · notation

Failures cluster. Banks fall in waves, power lines trip in cascades, stablecoins depeg on the same afternoon, cloud services go down together. Every industry has a word for it (contagion, systemic risk, cascading failure, correlated default), and most modelling mistakes in every one of those industries come from the same confusion: there are two completely different reasons failures arrive together, they demand opposite remedies, and almost every real disaster is a blend of both.

This article asks one question five times: why did these things fail together? First for systems that share a hidden common cause (§ 02), then for systems where failure genuinely spreads (§ 03), then at the boundary where small changes in coupling flip a system from forgiving to explosive (§ 04), then for real events, which mix the two (§ 05), and finally for portfolios of failures, where the whole is reliably not the sum of the parts (§ 06). § 07 is the honesty section: what it costs, in data, to measure any of this, and why the network diagram on the wall will not tell you.

Wthe exposure matrix: how much each unit holds of, or depends on, each channelΛimpact per channel: how much a unit of distress moves the thing everyone sharesVthe linkage matrix V = −WΛWᵀ: how much unit j is hurt when unit i is shockedN_effthe effective number of independent units in a portfolio of exposureskthe number of units failing togetherR(k)the sum-of-parts ratio: (sum of individual damages) ÷ (measured joint damage)
§ 01 / 07

The Two Definitions.

Shared exposure means everyone was secretly holding the same thing. Propagation means one failure changes the state of the survivors. The test: cut every channel between units and see if the clustering survives.

Start with the distinction the rest of the article lives on.

Definition 1.1: Shared exposure (correlation without propagation)def-1-1

Units fail together because a common shock hits an exposure they all hold. Nothing travels between them; each would have failed exactly the same way with the others absent. The heatwave reaches every cable; the rate shock reaches every bond portfolio; the bad training data reaches every model fine-tuned from it. Joint failure probability is set by the overlap of exposures, not by any interaction.

Definition 1.2: Propagation (true contagion)def-1-2

The failure of one unit changes the state of the others: a tripped line reloads its neighbours, a fire sale moves the price everyone else marks their book at, a failed counterparty becomes a loss on someone else’s balance sheet, a redemption drains the pool the next redeemer claims from. Joint failure probability depends on the dynamics of spread: remove the first failure and the second never happens.

The operational test that separates them: sever every channel between units (in a model, by intervention; in thought, honestly) and ask whether the clustering survives. Under shared exposure it survives untouched (the common cause is upstream of the units). Under propagation it vanishes.

The distinction matters because the remedies are opposite. Shared exposure is fixed by actual diversification (holding genuinely different things) and cannot be fixed by margin, because the shock arrives everywhere at once and margins are correlated too. Propagation is fixed by margin, buffers, and circuit breakers. It cannot be fixed by diversification, because the channel, not the holdings, carries the risk. Applying the wrong fix does not merely waste money; § 05 shows it often makes the other mechanism worse.

§ 02 / 07

Shared Exposure: Diversification You Don't Have.

Overlapping exposures make the effective number of independent units collapse toward one, and make independence assumptions fail exactly on the worst days.

If unit i holds exposure weights wiw_i over shared channels, the portfolio that looks like N names is really Neff=1/(θCθ)N_{\text{eff}} = 1/(\theta^\top C\,\theta) independent ones, where C is the correlation the overlaps manufacture. Two things about this number surprise people every time. First, it degrades fast: modest overlaps in the same few channels pull

NeffN_{\text{eff}} toward 1 long before the holdings look similar on paper. Second, it is invisible in calm data: shared exposure only expresses itself when the common channel moves, which is precisely when you needed the diversification to be real.

The quantitative signature: build the same system twice, once with the common shock shared, once with each unit given its own independent draw of the same marginal distribution, and compare tails. The independent version is not slightly optimistic; it is optimistic specifically about the events you built the model to see. In a heat-stressed island grid simulation, the correlated system exceeded the independent model’s “1-in-100” daily outcome 3.6% of the time; in a larger pre-registered grid study, 11.35%. The independence assumption was most wrong under exactly the common shock it ignored.

Fig. 01: Shared, or just similar? Illustration · you set the sharing

10 units, one shared shock, and a single dial for how much of it they hold in common. At the left end each unit draws its own day; at the right end they all share the same draw. Watch the two things that move and the one that does not: the joint distribution of the portfolio loss (petrol) fattens its tail as sharing rises, while the per-fund marginal (gold inset) is identical at every setting. The dashed line is the independent model's 1-in-100 day; the readout counts how often the correlated world clears it.

independent 1-in-100one fund: unchangedPORTFOLIO LOSS · BOND FUNDS
Shock sharing
0.60
Independent 1-in-100
3.42
Correlated breaches it
5.0%
vs nominal 1%
×5.0
Model: one common factor + idiosyncratic noise, identical marginals by construction · N = 10, seeded
The inset marginal never moves while the joint tail fattens: the diversification you appear to hold is undone by what you share, and it fails exactly on the shared shock's worst days.
§ 03 / 07

Propagation: The Reload Operator.

One quadratic form keeps appearing: damage flows through exposures, moves the shared channel, and comes back through everyone else's exposures.

Propagation needs a channel with impact: distress that moves something others are exposed to. The bookkeeping is the same across industries. Unit i is shocked and sheds RiR_i of its exposures; channel k moves by λkiRiwik\lambda_k \sum_i R_i w_{ik}; unit j, holding

wjkw_{jk} of that channel, marks the damage. The elasticity of *j*'s loss to a shock at *i* is:
Definition 3.1: The linkage matrixdef-3-1
V=WΛW\mathcal{V} = -\,W\,\Lambda\,W^{\top}

Exposures out, impact through the shared channel, exposures back. In fire-sale finance this is Greenwood–Landier–Thesmar’s vulnerable-banks operator; in a DC power grid the same shape appears as flow redistribution after an outage; in a supply chain it is capacity reallocation after a supplier fails. The physics differ; the algebra of “your problem becomes my problem, weighted by what we both touch” does not.

Iterating the operator gives rounds of contagion; whether the rounds die out or amplify is a spectral property of V\mathcal{V}, which § 04 turns into the only dial that matters.

Fig. 02: One engine, three skins Illustration · same engine, three renderers

One cascade engine, three costumes. Each node is one of banks holding overlapping books; each edge is a shared holding. Click any node to make it fail first. A neighbour then fails once enough of its neighbours have, and the wave runs until it stops. Flip the skin and the picture changes; the set of nodes that fall does not, because it is the engine's property, not the drawing's. Headroom lifts every node's buffer at once.

1GOLD RING = IGNITION · NUMBER = ROUND
Failed / total
2 / 11
Cascade share
18%
Rounds to settle
1
Headroom
1.00
Model: Watts threshold cascade on a seeded graph · click ignites a node, headroom lifts every buffer at once
Banks, grid, or supply chain: the picture changes, the set of nodes that fall does not. The cascade is a property of the coupling, not of the diagram drawn over it.
§ 04 / 07

The Phase Transition.

Coupling density is not a dial with proportional response. Below a threshold, failures stay local; above it, system-scale cascades appear abruptly. Event sizes go heavy-tailed.

Add coupling to a system one link at a time and watch the largest cascade it can produce. For a long stretch, almost nothing happens: a failure knocks out a neighbour or two and stops. Then, across a narrow band of coupling, the behaviour changes character: a single failure can now, with real probability, take down a system-sized fraction of the whole. Plot cascade size against coupling and you do not get a ramp; you get a step with a knee, and the knee is sharp.

Definition 4.1: The cascade thresholddef-4-1

Each failure triggers, on average, some number of further failures. Below a critical coupling density that number is less than one, and any cascade is a branching process that peters out: damage stays local. Above it the number exceeds one, and a cascade can grow without bound relative to the system. The crossing wears different names in different fields (the percolation point on a network, the epidemic threshold R0R_0 = 1 in contagion models, the instability line in Caccioli et al.’s overlapping-portfolio analysis), but it is one phenomenon: a branching ratio passing through one. The transition is abrupt in the size of the worst event, not the average one, which is exactly why averages give no warning of it.

Intuition 4.2: Efficiency and fragility are the same knobintu-4-2

Headroom (spare margin, spare capital, spare capacity, spare transmission rating) moves the knee: with more of it you can pack in more coupling before the system tips. But headroom never removes the knee; it only relocates it. And there is an uncomfortable corollary. A system run efficiently is, by definition, one operated with little idle headroom, which parks it close to its knee by construction. Efficiency and fragility are not opposing goals you dial between; near the threshold they are one number read from two directions. The most optimised version of a system is the one sitting nearest the cliff. The optimisation pressure is continuous, while the cliff is discovered all at once.

Near and above the knee, cascade sizes stop having a usable average. Most failures are still small, a few are system-scale, and the few dominate any total you care about. That is the heavy-tailed regime, where reasoning with means and standard deviations misleads in precisely the expensive direction. The apparatus for handling it honestly (what a tail exponent is, why averages mislead, how far you may extrapolate) is the whole subject of the extreme-value-theory background, which this section simply hands you off to. Here you need only the fact that the sizes are heavy, not the machinery for pricing them.

Fig. 03: The knee Illustration · branching-process model

Cascade sizes for a system whose branching ratio μ, the average number of further failures each failure triggers, you set through coupling and headroom. Left: the chance a cascade reaches size ≥ k (log-log). It falls away when μ < 1, straightens toward a line as μ approaches 1, and lifts onto a plateau once μ > 1. The plateau height is the probability of a system-scale event. Right: that same probability against coupling. It is not a ramp; it is flat, then a knee, and raising headroom slides the knee to the right without ever removing it.

11010011e-11e-21e-3system-scale plateauCASCADE SIZE ≥ k (LOG) · PROBABILITY (LOG)00.51knee μ=1audits assumeefficiency parks itCOUPLING DENSITY
Branching ratio μ
1.05
Regime
supercritical
P(system-scale)
9.0%
Knee at coupling
0.59
Model: total progeny of a Poisson(μ) branching process, μ = coupling × (1.9 − 0.7·headroom) · Borel–Tanner pmf, extinction fixed point
The right panel is the point: system-scale risk is flat, then a knee, not a ramp. Headroom relocates the knee; efficiency pressure walks the operating point toward it.
§ 05 / 07

Trigger × Amplifier.

Real disasters are a common shock that lights the match and a propagation channel that spreads the fire. Diagnosing only one half prescribes the wrong fix. Sometimes the fix for one worsens the other.

Almost no real disaster is purely one mechanism. The recurring template is a trigger (a shared shock that fails several units at once) followed by an amplifier (a propagation channel that turns those first failures into more). The shock lights the match; the channel spreads the fire. Reading only one half is the most common and most costly error in the subject, because the two halves take opposite remedies (§ 01) and the remedy for one routinely worsens the other.

Example 5.1: Three disasters, each a blendex-5-1
  • Heatwave into a grid. The heat is the shock: it derates every line at once and pushes several to their limits. Overload redistribution is the amplifier: a line that trips reloads its survivors, some past their limits. The heat alone would have stressed the grid; the channel is what turns stress into a blackout.
  • Rate shock into bond funds. A rate move marks down every fund holding the same duration: shared exposure, felt everywhere at once. Then forced selling drives the price down further and marks down everyone else, including funds that took no view and are coupled only through the asset they share. The shock became a channel, and the channel manufactured more shock.
  • A bank failure into a stablecoin. One bank fails: a localised event. But every holder of a coin exposed to that bank reprices in the same hour (shared exposure), and then redemption mechanics (the first redeemers draining the reserve the next redeemers claim from) propagate the repricing into a run. One trigger, two amplifiers stacked on it.
Pitfall 5.2: Fixing the mechanism you don't havepit-5-2

Cross what is really happening against what you assume is happening. The two matched cases are fine: diversify against shared exposure, add margin and breakers against propagation. The two mismatched cases are where the losses live.

  • Margin against a common shock. You assumed propagation and posted collateral. But a shared shock calls every position on the same day, and the collateral (being correlated too) is impaired on that same day. The buffer evaporates exactly when it is needed.
  • Diversification against a channel. You assumed shared exposure and spread across more counterparties. But under propagation each new counterparty is a new link, so the fix enlarges the surface the failure travels across. “Diversifying” here increases the very risk it was bought to reduce.

Neither error merely wastes money; each tends to load the other barrel.

You cannot diversify away a common shock, and you cannot hold margin against a channel you haven't mapped.

§ 06 / 07

Joint Failures Are Not Additive.

The damage of k failures is reliably not k times the damage of one. The sign of the gap is an empirical property of the system, not an assumption you get to make.

Ask what k simultaneous failures cost. The tempting answer is k times the cost of one, and it is almost never right. Whether the true cost runs above or below additive is not a modelling choice you get to make; it is an empirical property of the specific system. It can change sign as k grows.

Definition 6.1: The sum-of-parts ratiodef-6-1

Take a set S of k units. Fail each one on its own and measure the damage DiD_i; then fail the whole set together and measure the joint damage D(S)D(S). Define

R(k)=(iSDi)/D(S)R(k) = \big(\textstyle\sum_{i\in S} D_i\big)\big/\,D(S)

the sum of the parts over the measured whole. R(k)=1R(k) = 1 exactly when the failures do not interact: the parts add. R(k)>1R(k) > 1 means the whole is less than the sum (sub-additive: the failures partly cover for one another). R(k)<1R(k) < 1 means the whole is more than the sum (super-additive: the failures compound). The ratio is measured, not assumed, and its being different from one is the interaction.

Example 6.2: Both signs happenex-6-2

Grids tend to come out sub-additive (R(k)>1R(k) > 1). The first outage already sheds the load the second one would have stranded; by the time the second line fails, there is simply less left to lose. It is a saturation effect, and it makes joint failures less than additive. Fire-sale systems tend to come out super-additive (R(k)<1R(k) < 1). Each forced sale deepens the price impact the next sale sells into, so damages compound and the joint loss overshoots the sum. Same question, opposite sign. The sign is a fact about the mechanism (saturation versus compounding), readable off measurement, not off intuition.

Pitfall 6.3: N−1 planning meets a curved worldpit-6-3

Reliability rules built on N−1 (survive any single failure) plus a silent additivity assumption misprice N−k events in whichever direction the system actually bends: over-cautious where it saturates, dangerously optimistic where it compounds. Worse, the direction can flip inside one system: compounding among the first few failures (small k), saturation once enough has already broken (large k). A single measured R(k)R(k) curve is therefore a claim about one system over one range of k, never a constant you may carry elsewhere.

Assumption. R(k)R(k) is measured on the same system and the same stress distribution it will be used to reason about; a ratio found under one regime does not transfer to another.

Fig. 04: Sum of parts Illustration · toy engine

The cost of failing k units together, divided by the cost of failing them one at a time. The dashed line at 1 is the world where failures do not interact: the parts add. Above it (petrol, a saturating engine) the whole is less than the sum: the first failures shed the load the later ones would have stranded. Below it (gold, a compounding engine) the whole is more: each failure deepens the next. Solid curves take the worst units first; the faint dots are random sets, so you can see how far the typical case sits from the corner.

0.60.81.01.21.41.61.82.02.22.42.616121824R = 1 · no interactionsub-additivesuper-additiveNUMBER OF UNITS FAILING TOGETHER · k
Saturating R(24)
2.66
Compounding R(24)
0.62
Interaction strength
0.55
Sign of the gap
both, opposite
Model: 24 heterogeneous units · saturating (capacity-bounded) vs compounding (quadratic cross-term) joint damage · worst-case curves, random-set scatter
Both signs are real and both are measured, not assumed. The gap between the solid worst-case curve and the cloud of random sets is why a single R(k) needs its set said out loud.
§ 07 / 07

What This Costs to Measure.

Influence is not readable off the network diagram, per-pair contagion estimates need brutal data budgets, and the honest unit of reporting is a paired contrast with its interval. These are claims this article states and companion pieces defend.

The sections above describe structure. This one is about what it costs to measure that structure on a real system. The answer is sobering enough to change how the numbers should be reported. Three findings, each stated with its pointer and left to be defended in full elsewhere rather than developed here.

The diagram lies. The wiring picture on the wall (who is connected to whom) is a poor guide to who is dangerous. In a pre-registered power-grid study, the degree of a link’s endpoints explained almost none of its measured interventional influence: rank correlation ρ ≈ 0.18–0.22, against a pre-registered 0.30 bar for “worth using”. The only cheap topological proxy that carried real signal was thermal rating, and it captured essentially all of the reproducible part. The lesson generalises past grids: danger rankings are a property of the dynamics, not of the picture, so they have to be simulated or intervened for, system by system.

Budgets bite. Estimating a single per-pair contagion probability well takes more stressed data than intuition expects. In the same study, per-pair estimates from 8 stressed weeks had Brier skill indistinguishable from zero (no better than quoting the base rate), while at 32 weeks skill rose to +0.40. Between those budgets the only content that transferred was the overall crossing rate, not any pairwise structure. A contagion matrix quoted without the budget it was estimated on is noise with a decimal point. [→ forthcoming background: simulation budgets & paired contrasts]

Ceilings before victory laps. A claim of the form “our model recovers the true influence ranking at ρ = 0.77” is empty until the reliability of the truth has itself been measured against the same perturbation the claim faces. If the ground-truth ranking only agrees with itself at ρ ≈ 0.57 across independent redraws of the noise, then 0.77 is not skill; it is impossible, an artefact of scoring against one lucky draw of the truth. Measure the ceiling before admiring the height. [→ forthcoming background: measurement ceilings]

What survives all three cautions is a modest reporting discipline. The honest unit is not a network diagram, and not a single contagion number, but a paired contrast with its interval: this configuration against that one, same shocks, same seeds, the difference reported with the uncertainty the measurement budget can actually support. It is the same discipline the calibration background asks for from the other side: state the uncertainty, or you have not stated the result.

You should now be able to take any story about things failing together and take it apart. Ask what share was shared exposure and what share was propagation, and test it by cutting the channels in thought (§ 01). Ask where the system sits relative to its knee, and which way its headroom is moving it (§ 04). Ask the sign of R(k)R(k), and whether it flips (§ 06). And ask what budget the contagion numbers were estimated on before believing any of them (§ 07). The two mechanisms are simple; almost every expensive mistake is naming the wrong one.

References the shelf behind the article
  1. Greenwood, R., Landier, A., Thesmar, D. (2015). Vulnerable banks. J. Financial Economics. The linkage operator V = −WΛWᵀ of § 03.
  2. Caccioli, F., Shrestha, M., Moore, C., Farmer, J.D. (2014). Stability analysis of financial contagion due to overlapping portfolios. J. Banking & Finance. The overlapping-portfolio phase transition of § 04.
  3. Gai, P., Kapadia, S. (2010). Contagion in financial networks. Proc. R. Soc. A. The propagation cascade and its tipping condition.
  4. Glasserman, P., Young, H.P. (2016). Contagion in financial networks. J. Economic Literature. A survey weighing how much network propagation actually adds over shared exposure.
  5. Elliott, M., Golub, B., Jackson, M.O. (2014). Financial networks and contagion. Amer. Econ. Rev. Integration versus diversification, and why more links cut both ways (§ 05).
  6. May, R.M., Levin, S.A., Sugihara, G. (2008). Ecology for bankers. Nature. Systemic risk read across ecology and finance; the two-mechanism split is not unique to markets.
  7. Watts, D.J. (2002). A simple model of global cascades on random networks. PNAS. The cascade threshold and the sharpness of the knee (§ 04).
  8. Dobson, I., Carreras, B.A., Lynch, V.E., Newman, D.E. (2007). Complex systems analysis of series of blackouts. Chaos. Grid cascades, load saturation, and heavy-tailed blackout sizes (§§ 04, 06).
  9. Battiston, S., Puliga, M., Kaushik, R., Tasca, P., Caldarelli, G. (2012). DebtRank: too central to fail? Scientific Reports. Interventional influence versus topological centrality (§ 07).