Loading…
Loading…
227 on-domain prediction-market questions tracked · 198 resolved · 29 open · 4 map to a scored material in our register.
Rung 1 of the calibrated forecasting layer. It ingests the free public prediction-market read APIs and keeps only the questions on our axis — tariff, sanctions, export-control and trade-policy events (never diplomatic / back-channel questions). Where a question names a critical material, it is mapped to the scored-material vocabulary so it lines up with the register's supply-side view. The public market price is a calibration benchmark: a continuously-scored, copy-proof track record is the moat the forecasting layer accrues.
Dual-score discipline. A market-implied probability is an alternative / OSINT signal — shown here in its own track, never merged into the Tier-1 official register or the policy-pressure number. The divergence between the market, our model and the official legislative stage is itself the signal. Probabilities below are real public market prices, not model outputs — a calibrated model probability (Rung 2) ships only once this harness can prove it is calibrated.
Sibling channel: self-calibration backtest — scores our own forward axis-2 rubric against our own register (no external source).
Scored over 240 genuine forecast-time points — prices recorded while the market was open, then scored against the eventual resolution. Lower is better for both metrics.
Provenance. 217 points backfilled from already-resolved markets' open-time price history — two free public sources (Polymarket CLOB price history + Kalshi daily candlesticks), sampled by the SAME rule at the midpoint of each market's traded life (never the settled price); 23 from live daily snapshots. Backfill includes every on-domain resolved market with recoverable history — no cherry-picking; markets whose price history is empty are skipped as an honest coverage gap, never imputed. Equity / earnings markets (e.g. “will <chipmaker> beat quarterly earnings?”, admitted by the “semiconductor” keyword) are now excluded — they are company-performance bets, not trade-policy events, so the board scores genuine statecraft questions only.
R43 proved the free market is calibrated on our domain — but a well-calibrated free price argues against paying us. The licensing thesis rests on the opposite: that our hand-curated, git-versioned register conditions the outcome beyond what the market price already encodes. For each resolved market that maps to a register material, we derive an ordinal register signal as-of the forecast date (using only actions announced on/before it — no lookahead) and cross-tab it against the resolution, with the market's own Brier on the same subset alongside it.
As-of-D signal rule (fixed, pre-registered). register actions targeting the market's mapped material with announced_date ≤ D, counted in the 90 days on/before D (C90): active-escalation if C90 ≥ 3, precedent if C90 in 1–2, none if C90 = 0. Eligibility. a genuine material-trade-policy event: (1) not an equity/earnings market, (2) the mapped material appears as a whole word (not a substring), (3) the question carries a trade-policy verb (export/tariff/sanction/ban/…).
The edge cross-tab needs 10 genuine material-trade-policy resolved markets; only 5 of 13 mapped rows qualify. The rest are pollution in the coarse keyword detector we refuse to score on: 0 equity / earnings markets and 8 substring mis-maps. The binding constraint is no longer pollution but genuine corpus thinness: the Kalshi candlestick backfill added 49 resolved trade-policy points to the calibration board above, but they are country-level tariff-rate brackets and trade-deal questions with no specific material reference, so they widen the market-calibration board without widening the edge set. The material-specific questions that would (copper / GPU-export / lithium tariffs) are still open mid-2026 — the resolved material-trade-policy corpus public markets have priced so far is dominated by two clusters (US semiconductor tariffs, China rare-earth restrictions). That is itself the honest finding: proving register edge from free public markets alone likely has to wait for those material-specific markets to settle.
Free public sources only (Polymarket / Kalshi read APIs; Metaculus needs an authenticated token, so it is disabled rather than faked). On-axis filtering keeps only trade / sanction / export-control / trade-policy questions. This is the no-regret Rung-1 foundation of the forecasting layer — it measures whether we are calibrated; it does not yet publish a probability.