Causal-chain visualizer — reader-test kit
This is the falsification harness specified by chunk 3 of docs/strategy/projects/causal-chain-visualizer.md. Gate 5 of docs/strategy/gates.md flips from building to passing when the test result recorded at the bottom of this file shows ≥4 of 6 readings identifying all three target facts on the regime they were shown.
The kit is fully self-contained. Hand a non-IPTM-literate reader the URL of one regime page plus the Reader instructions below. Do not let them read any other MacroLens prose, this file included. After they submit, the administrator records the result in the Result log at the foot of this file.
Why this test, in one paragraph
The product claim Gate 5 carries is that the causal-chain visualizer makes a multi-action policy cascade legible without the reader reading any explanatory prose — they can read the shape alone and recover (a) which action started the cascade, (b) who responded to whom, (c) the direction of the reflex (e.g., "China responded to the US"). If a reader cannot do that, the visualizer has failed the product test; it does not matter how rigorous the underlying data is. Sending the SVG to a non-IPTM-literate reader is the only way to test this — a reviewer who has read the case study already knows the answers. The point of running two readers per regime is to detect ambiguity that a single reader's idiosyncrasy might mask.
Reader profile we need
- Numerate but non-specialist. Comfortable reading a chart; does
not need to know what BIS, MOFCOM, CBAM, or "responds_to" means.
- No prior MacroLens exposure. A friend or colleague who has not
used the platform. (If they have, ask a different person.)
- English-comfortable. The labels on the SVG are English.
- 20–30 minutes of attention. That is enough for one regime.
We need 2 readers per regime × 3 regimes = 6 readings. The same person may read more than one regime if the regimes are administered on different days (otherwise their second reading is biased by their first). Recommended: 2 different people per regime, total 4–6 distinct humans.
Reader instructions (paste this verbatim to the reader)
> I am running a test on a chart we built. There is no right answer I > need you to find; I need to know whether a chart someone has never > seen before can be read by itself. Please do these three things in > order, and do not read any other text on the page — just the > chart and its in-chart labels. > > 1. Open: <URL> (one of the three regime URLs in the table below). > 2. Look at the dots and arrows on the chart. The dots are > government actions. The arrows mean one action was a response to > another action. The number on each arrow is the number of days > between them. > 3. Answer these three questions (in your own words; one sentence > each is enough): > - Q1 — foundation: Which action started this chain? (Point > to one dot.) > - Q2 — response: Pick one arrow. Who responded to whom? > (E.g., "Country X responded to Country Y" or "the second dot > on the right responded to the first dot on the left".) > - Q3 — reflex direction: Across the whole chart, which > country is reacting, and which country is being reacted to? > (E.g., "China reacted to the US" or "Vietnam reacted to the > EU" or "they reacted to each other".)
Send the reader the link from the Regime URL column for whichever regime you have allocated them. Do not tell them the regime title.
The three regimes
| ID | Regime URL (production) | Regime URL (local dev) | What the SVG shows |
|---|---|---|---|
| A | https://macro.ciconialabs.com/regime/2024-china-ga-ge-retaliation-cycle | http://localhost:3010/regime/2024-china-ga-ge-retaliation-cycle | 5 actions in 2 swimlanes (US BIS + China MOFCOM), tit-for-tat retaliation cycle with the reflex collapsing from 269 days to 1 day. |
| C | https://macro.ciconialabs.com/regime/2026-eu-cbam-policy-diffusion | http://localhost:3010/regime/2026-eu-cbam-policy-diffusion | 5 actions in 3 swimlanes (EU + Vietnam + UK), single-anchor regulatory cascade with policy-architecture export to two third countries. |
| E | https://macro.ciconialabs.com/regime/2026-eu-russia-forking-architecture | http://localhost:3010/regime/2026-eu-russia-forking-architecture | 6 actions in 2 swimlanes (US/G7 + EU), forking sanctions architecture with a legal-base pivot at action 5. |
If the production URL is not reachable (the site is currently private), use the local-dev URL after the administrator runs npm run dev on the VPS or paths to the dev port via SSH tunnel. The SVG is identical between the two.
Scoring rubric (administrator-only — do not share with reader)
Each reading produces three answers. Score each answer PASS if it matches the grader's key below, FAIL if it doesn't. A reading is a clean read if all three answers PASS.
The gate flips to passing when:
- At least 4 of the 6 readings are clean reads, AND
- At least 1 clean read per regime (no regime sweeps 0/2).
If the test fails, the failure mode determines next action per the project plan:
- Same question fails across multiple regimes → presentation issue
shared by the component (e.g., the arrow head is unclear). File a refinement spec back to wake-graph.
- Different questions fail in different regimes → likely a
per-regime data issue (action labels too long, anchor not clear). File per-regime fixes.
- One regime fails entirely while the other two pass cleanly →
the failing regime's chain is poorly shaped for V1; drop it from the V1 set and rerun. The Gate-5 target is ≥3 regimes, so two passing + one dropped + one promoted from the V2 backlog (Indonesia hilirisasi or US chip-control) is acceptable.
Grader's key (administrator-only — do not share with reader)
Regime A — Ga/Ge retaliation cycle
- Q1 (foundation): the **2022-10-07 US BIS advanced AI chip
controls** dot. It is the top-left dot on the chart (earliest date, left swimlane). A reader who points to the 2024-12-03 MOFCOM dot (the anchor, which is rendered larger) instead is FAIL — anchor size ≠ first cause. Watch for this confusion in the scoring.
- Q2 (response): any arrow that correctly identifies the
responder as the lower dot and the trigger as the higher dot (the visualizer convention is responder→trigger, arrowhead on the trigger). Accept any of the four edges in this regime. Reader does not need to name the action_type; identifying "the second dot responded to the first dot" is sufficient.
- Q3 (reflex direction): "China reacted to the US" OR "the US and
China reacted to each other" are both PASS. Tit-for-tat is the thesis; either rendering of it clears. "The US reacted to China" is FAIL — the chain originates with the US BIS anchor.
Regime C — CBAM diffusion
- Q1 (foundation): the 2026-01-01 CBAM definitive phase dot
(left swimlane, EU; rendered as the anchor with +4px halo). It is the only EU dot with explicit responses pointing back at it from Vietnam and the UK. PASS for "the CBAM dot" / "the EU 2026 dot" / "the big EU one in the middle of the chart". Pointing to the 2025 Clean Industrial Deal or Steel Action Plan is partial PASS — those are forward-signal precursors but not the responded-to anchor in the responds_to edges. Score as FAIL if explicitly named instead of the CBAM anchor.
- Q2 (response): either the Vietnam 2026-01-19 Decree 29 → CBAM
edge, or the UK 2026-03-18 Finance Act → CBAM edge. PASS for "the Vietnam dot responded to the EU dot" or "the UK dot responded to the EU dot".
- Q3 (reflex direction): "Vietnam and the UK reacted to the EU"
is the clean PASS. "Vietnam reacted to the EU" or "the UK reacted to the EU" alone is also PASS (reader saw one of two parallel responses). FAIL only if the reader has the arrow direction wrong (e.g., "the EU reacted to Vietnam").
Regime E — EU-Russia forking architecture
- Q1 (foundation): the 2022-12-05 G7+EU oil price cap dot —
top-left of the chart, US/G7 swimlane, earliest date. PASS for pointing to it explicitly. Partial PASS for the 2024-06-24 EU 14th-package dot if the reader read the chart left-to-right and identified the first EU-swimlane dot (this is honest confusion — the regime has two swimlanes and the foundation is in the US/G7 swimlane, not the EU swimlane). Score partial PASS as FAIL for Q1-PASS purposes but flag in the result log — repeated partial PASS across readers means the foundation is genuinely ambiguous and we need to label the anchor more clearly.
- Q2 (response): any edge from an EU dot back to the G7 oil price
cap, OR any edge within the EU swimlane (e.g., the EU 20th package → EU 19th package edge). PASS for "the EU [dot N] responded to the US [dot 1]" or "the EU [later dot] responded to the EU [earlier dot]". The within-EU response is correct — the EU 18th/19th/20th packages explicitly close circumvention rails opened by the previous package.
- Q3 (reflex direction): "the EU is reacting to the US oil price
cap, then to itself as Russia adapts" is the clean PASS. PASS for any answer that captures (a) the EU is the responder and (b) the trigger is either the US/G7 anchor or prior EU actions. The forking-architecture thesis is more nuanced than Q3 can fully test, so the bar for Q3 here is "the reader has the responder/trigger direction right", not "the reader has restated the thesis".
Distribution & administration checklist
Before sending the URL to a reader, run through this:
- [ ] Confirm the regime page renders cleanly (open the URL on the
administrator's own machine, check the SVG is visible, the dots and arrows are drawn, no console errors).
- [ ] Confirm no "DATA ERROR — TRIAGE" red arrows (the visualizer
draws these when responds_to resolves to a date earlier than the trigger, which should never happen but is guarded for). If one appears, file an audit issue with wake-graph before running the test — the SVG is showing a data-integrity bug, not a presentation problem.
- [ ] Confirm the reader has not previously seen MacroLens. (Ask:
"Have you ever opened macro.ciconialabs.com?")
- [ ] Send the reader only the URL and the Reader instructions
block above. Do not send this file. Do not send the regime title. Do not send the case-study URL.
- [ ] Record the reader's three answers verbatim in the Result log
below. Verbatim matters — paraphrasing erodes the scoring rigour.
- [ ] After the reading, optionally show the reader the case-study
page (docs/intelligence/cases/<slug>.md) and ask whether their reading was correct. Their reaction (surprise / confirmation / confusion) is qualitative signal worth recording in the result log even though it does not affect the scoring.
Result log
| Reading # | Regime | Reader (initials or alias) | Date | Q1 answer | Q1 score | Q2 answer | Q2 score | Q3 answer | Q3 score | Clean read? |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | A | _pending_ | ||||||||
| 2 | A | _pending_ | ||||||||
| 3 | C | _pending_ | ||||||||
| 4 | C | _pending_ | ||||||||
| 5 | E | _pending_ | ||||||||
| 6 | E | _pending_ |
Aggregate: _N_ of 6 clean reads; per-regime breakdown _A=?/2, C=?/2, E=?/2_. Gate-flip threshold: ≥4 clean reads AND ≥1 clean read per regime.
Verdict: _pending administration_
Caveats and known limits
- **The administrator and the grader are the same person in
practice*, which is a structural weakness — the grader knows the thesis and may be unconsciously lenient. Mitigation: write each reader's verbatim answer in the result log before* looking at the grader's key, then score in a second pass.
- The grader's key for Regime E is softer than for A and C. The
forking-architecture thesis has more moving parts than the simpler retaliation-cycle and diffusion-cascade theses, and Q3 there rewards "directional correctness" rather than "thesis recovery". Two clean reads on E mean "the SVG conveys direction"; they do not mean "the SVG conveys the legal-base pivot". The latter requires prose and is correctly out of scope for V1.
- **A reader who is friends with the administrator may be charitable
in framing their answer.** When in doubt, score the literal words the reader produced, not the most-charitable reading.
- **A pass result here is necessary but not sufficient for the
product claim.** Gate 5 testing the visualizer does not test whether the underlying case studies are correct — Gate 1 carries that. A clean reader-test pass on a regime with a wrong thesis would still pass Gate 5 but invalidate Gate 1.
- The V2 regime backlog (Indonesia hilirisasi, US chip control)
is held out of this test. Those regimes have V1-incompatible shapes per the design doc — running the same reader test on them before they are V2-rendered would underestimate the visualizer's ceiling.
Where this kit fits in the project plan
docs/strategy/projects/causal-chain-visualizer.md chunk-3 line: "Reader-test pass. Run the falsification harness on all 3 regimes. If any fail, file a refinement spec back to wake-graph. If all pass, flip Gate 5 `state:` to `passing`."
This file IS the falsification harness. The administrator runs the test, records the results, and a future strategy tick reads the result log and flips Gate 5 if the criteria are met. If the test fails, the strategy tick files a refinement spec to wake-graph (an ops/queue/graph.md append) before retrying.
The kit's design tries to honour Gate 5's spirit: a non-IPTM-literate reader is the only honest reviewer for a "reads without prose" claim. A self-administered simulation by the wake itself would replicate the problem of the same head writing the SVG and grading it — exactly the failure mode the harness exists to detect.