Does the supply-risk score actually work? An honest backtest
Every material dossier in this atlas carries a 0-100 supply-risk score. This piece asks the question a reader should ask before trusting it: does the score anticipate which materials actually get weaponized, or does it just look plausible?
The honest test, not the flattering one
The score has five factors: producer concentration, EU import reliance, substitutability, policy pressure, and price stress. Policy pressure is built partly from the same export-control register we'd use as "ground truth" — so testing the full score against that register proves nothing; it's circular by construction.
The fair test strips policy pressure out and asks whether the structural score alone — concentration, import reliance, substitutability, all knowable before any government acts — lines up with what has historically been weaponized.
Two ways of defining "weaponized," on purpose
A single hand-picked list of famous cases (the 2010 China rare-earth embargo, the 2023-25 Ga/Ge/Sb/W/Te/Bi controls, the 2020 Indonesia nickel ban) is an easy trap: the list-builder unconsciously picks cases that fit the story. So this backtest runs two panels and reports both:
1. Systematic (primary) — an objective, mechanical rule: a material counts as weaponized if one of its own top-2 producing/refining countries issued a restrictive export-control action (severity ≥ 4) against it, per our own register. Every material gets run through the identical test — no curation. This turns up 36 qualifying actions across 16 materials, including cases (uranium, lithium, phosphate) the famous-cases list misses entirely. 2. Flagship (secondary) — the 8 famous named cases, kept only as readable case-study color and a sanity check. The two panels agree on 28 of 31 materials.
The result
| Panel | Structural AUC | What that means |
|---|---|---|
| Systematic (primary) | 0.76 | Meaningfully better than a coin flip (0.50), well short of perfect (1.00) |
| Flagship (secondary) | 0.82 | Similar signal, smaller sample |
Materials the structural score ranks in its top 10 are hit roughly 80% of the time. It is a real, useful prior — not noise. It is also not a crystal ball: it ranked manganese, magnesium, and niobium as high-risk with no control materializing (yet), and it ranked nickel and bismuth low despite both getting hit.
What this doesn't prove
This is a discrimination test on today's data against historical events, not a frozen-in-time forecast. We don't have 2009-vintage concentration data, so we can't say the score, computed in 2009, would have called the 2010 embargo. It answers "does current structural risk line up with what has historically been weaponized," not "would this have predicted the future." Read the score as an evidence-weighted prior worth factoring into a decision, not a guarantee.
We publish this the same way we publish the retired v3b composite's -11.2pp/yr backtest failure and the IPTM signal's null result: what doesn't fully work stays visible, next to what does.