How PEPIndex is calibrated against historical events and external benchmarks
PEPIndex is calibrated against historical events with well-documented outcomes. Each calibration
event carries a pre-calculated composite score and a documented "key gap" between what was
announced and what was applied. Resolved calibration events are locked, so their hand-verified
values never drift as the live pipeline runs.
Calibration events
Seed events with documented outcomes and pre-calculated composite scores. Authority identifies the legal basis for trade-domain events.
Event
Domain
Authority
Score
Outcome
Key gap
Section 232 Steel/Aluminum Global (2018)
Trade
Section 232
34.0
Partial reversal
25% global → ally exemptions (~45% scope)
Section 301 China Trade War / Phase One (2018–2020)
ECFR force-threat follow-through. An analysis of presidential military-force threats found a 9.1% follow-through rate (2 of 22 tracked) — the empirical anchor for the foreign-policy domain's high baseline reversal expectation.
Retail price pass-through. Published peer-reviewed estimates of retail pass-through for the 2025 tariffs (Cavallo, Llamas & Vazquez, 2026) are used to validate the economic-impact calculations for trade-domain events.
First-term trade-war rate references. Source for first-term calibration events (TACO-2018-*): dates, rates, and Phase One deal terms.
Yale Budget Lab Tariff Tracker. Independent cross-check of the import-weighted effective tariff rate and PCE goods/durables pass-through ranges.
Federal Register Proclamation 11021 (FR 2026-06960; 91 FR 18201) and its June 2026 amendment — source of record for the 2026 Section 232 metals events.
Known limitations
Classification subjectivity. Some events involve judgment about whether a statement is a genuine threat or rhetorical posturing; inter-rater testing mitigates but does not eliminate this.
Outcome timing. "Resolution" can be ambiguous for quietly-abandoned threats. A bounded monitoring window settles it, anchored on the stated deadline where the threat named one and on the threat date where it did not. At that edge, non-trade lapses finalize to indeterminate (excluded from the chicken-out rate).
Causal attribution. Post-announcement price moves may reflect concurrent factors; the four-group event-study design controls for trend but causal identification is imperfect.
Non-trade economic impact. Quantified economic impact is currently limited to trade-domain events; foreign-policy and domestic events are scored behaviorally only.
Political neutrality. A high PEPI score means a threat was not followed through — it makes no normative judgment about whether follow-through would have been desirable.
Economic inputs are latest-vintage, not point-in-time. Government statistical series are revised after publication, on no schedule we assume — benchmark and seasonal-factor revisions reach further back than the routine one, and the release calendar itself can slip. PEPI re-reads the full history each run, so a published economic figure reflects the currently published value for its reference period, not the value that existed when we first read it. A figure quoted from an earlier snapshot may therefore differ from the same figure today. The scope is bounded and verified rather than assumed: revisions move the price-impact measures and per-event economic columns, and move no composite score and no daily index value — the scorer's economic inputs are USITC tariff rates, not price statistics. See data vintage & revisions.
Coverage limits — read before citing a figure
Three published series are legitimately thin or structurally quiet. They are
correct as computed, but each is easy to misread, so we state them plainly rather
than let a number be quoted out of context.
The congressional reaction sub-index reads near zero, and that is structural.
It is built from legislative actions on bills the tracked threats reference — and
presidential threats rarely cite a bill ("I will impose a 100% tariff…" names no
legislation), so few tracked events ever acquire a legislative reference. The
measure is also a trailing activity window, so bill actions that do exist stop
contributing once they age out. There is deliberately no fallback that would
manufacture a non-zero value. Read it as "how often tracked threats intersect
the legislation they cite, recently," which is near zero. It is
not a measure of congressional activity in general. Subject-matter
matching — which would link a tariff threat to a trade bill it never names — is
deliberately unbuilt: its false-positive rate would put noise into a published
series.
The rhetorical domain is three calibration benchmarks, not a live sample.
Every event currently in that domain is a hand-curated, documented benchmark
(broadcast-license threats, "open up the libel laws," the subsidies/contracts feud
threat). Its headline figures — including a 100% reversal rate — describe
n=3 deliberately-selected historical cases, and must not be quoted as a
measured population rate. The other three domains carry live events.
Rhetorical Dilution is live but thinly measured — and its weight
does not mean its influence. The dilution dimension is measured only where
a follow-up statement could actually be matched to its event, which today is a
small minority of the catalogue. Where no statement matched, the input is
null and the dimension contributes nothing — we do
not assume “no softening occurred,” because absence of a matched
statement is not evidence of an unwavering threat. The same holds for the
victory-claim half of Framing Shift. Scope Reduction, by contrast, is measured on
the large majority of resolved events, because the resolvers judge it on evidence
they already read. So a composite score is best read as the dimensions that
could be measured for that event, not as all six carrying their nominal
weight. This is deliberately conservative: an unmeasured dimension never inflates
a reversal score.
The judicial sub-index measures activity, not influence. It is a
saturating measure of how actively the courts are engaging with matters we track —
filings activity-weighted by court, plus court-evidenced resolutions. A high value
means the courts are busy, not that they are reversing the
president; the separately-reported resolution count is the influence signal, and
it is frequently zero over a given window.
The principles behind the six dimensions, the sources and the validation posture are in the
📄 methodology overview; version history is on the
changelog.
Per-event provenance (Federal Register links, statutory authority, court rulings, tariff-rate
snapshots) is available through the API.