How PEPIndex is calibrated against historical events and external benchmarks
PEPIndex is calibrated against historical events with well-documented outcomes. Each calibration
event carries a pre-calculated composite score and a documented "key gap" between what was
announced and what was applied. Resolved calibration events are calibration_locked so
their hand-verified values never drift as the live pipeline runs.
Calibration events
Seed events with documented outcomes and pre-calculated composite scores. Authority identifies the legal basis for trade-domain events.
Event
Domain
Authority
Score
Outcome
Key gap
Section 232 Steel/Aluminum Global (2018)
Trade
Section 232
34.0
Partial reversal
25% global → ally exemptions (~45% scope)
Section 301 China Trade War / Phase One (2018–2020)
ECFR force-threat follow-through. An analysis of presidential military-force threats found a 9.1% follow-through rate (2 of 22 tracked) — the empirical anchor for the foreign-policy domain's high baseline reversal expectation.
Retail price pass-through. The 24% retail pass-through rate for the 2025 tariffs is used to validate the economic-impact calculations for trade-domain events.
First-term trade-war rate references. Source for first-term calibration events (TACO-2018-*): dates, rates, and Phase One deal terms.
Yale Budget Lab Tariff Tracker. Independent cross-check of the import-weighted effective tariff rate and PCE goods/durables pass-through ranges.
Federal Register Proclamation 11021 (FR 2026-06960; 91 FR 18201) and its June 2026 amendment — source of record for the 2026 Section 232 metals events.
Known limitations
Classification subjectivity. Some events involve judgment about whether a statement is a genuine threat or rhetorical posturing; inter-rater testing mitigates but does not eliminate this.
Outcome timing. "Resolution" can be ambiguous for quietly-abandoned threats. Two monitoring windows bound it, matched to whether the threat named a date: a deadline-bearing event stays live until its stated deadline + 90 days, a deadline-less event until its threat date + 270 days. At that edge, non-trade lapses finalize to indeterminate (excluded from the chicken-out rate).
Causal attribution. Post-announcement price moves may reflect concurrent factors; the four-group event-study design controls for trend but causal identification is imperfect.
Non-trade economic impact. Quantified economic impact is currently limited to trade-domain events; foreign-policy and domestic events are scored behaviorally only.
Political neutrality. A high PEPI score means a threat was not followed through — it makes no normative judgment about whether follow-through would have been desirable.
Economic inputs are latest-vintage, not point-in-time. Government statistical series are revised after publication, on no schedule we assume — benchmark and seasonal-factor revisions reach further back than the routine one, and the release calendar itself can slip. PEPI re-reads the full history each run, so a published economic figure reflects the currently published value for its reference period, not the value that existed when we first read it. A figure quoted from an earlier snapshot may therefore differ from the same figure today. The scope is bounded and verified rather than assumed: revisions move the price-impact measures and per-event economic columns, and move no composite score and no daily index value — the scorer's economic inputs are USITC tariff rates, not price statistics. See data vintage & revisions.
Coverage limits — read before citing a figure
Three published series are legitimately thin or structurally quiet. They are
correct as computed, but each is easy to misread, so we state them plainly rather
than let a number be quoted out of context.
The congressional reaction sub-index reads 0.0, and that is structural.
An event is linked to legislation only when its text either matches a curated
bill-alias list or carries an explicit bill citation ("H.R. 22",
"S.J.Res. 7") — and presidential threats almost never name a bill ("I will impose
a 100% tariff…" cites no legislation). Two things therefore hold the series at
zero, and both are real: few tracked events ever acquire a legislative reference,
and the measure is a trailing 30-day activity window, so bill
actions that do exist stop contributing once they age out. There is deliberately
no fallback that would manufacture a non-zero value. Read it as "how often
tracked threats intersect cited or aliased legislation, recently," which is
near zero. It is not a measure of congressional activity in
general. Subject-matter matching — which would link a tariff threat to a trade
bill it never names — is deliberately unbuilt: its false-positive rate would put
noise into a published series.
The rhetorical domain is three calibration benchmarks, not a live sample.
Every event currently in that domain is a hand-curated, documented benchmark
(broadcast-license threats, "open up the libel laws," the subsidies/contracts feud
threat). Its headline figures — including a 100% reversal rate — describe
n=3 deliberately-selected historical cases, and must not be quoted as a
measured population rate. The other three domains carry live events.
The judicial sub-index measures activity, not influence. It is a
saturating measure of how actively the courts are engaging with matters we track —
filings weighted by court tier, plus court-evidenced resolutions. A high value
means the courts are busy, not that they are reversing the
president; the separately-reported resolution count is the influence signal, and
it is frequently zero over a given 30-day window.
Full source-by-source lineage, scoring examples, and version history are in the
📄 methodology whitepaper.
Per-event provenance (Federal Register links, statutory authority, court rulings, tariff-rate
snapshots) is available through the API.