Consequences Dataset — Methodology
Phase 1 of the “Consequences” section: a dataset of U.S. federal political scandals since 1970, each scored on a documented consequence scale with cited receipts. This page documents how the dataset was built, the rubric applied to every entry, and every exclusion.
Working thesis (pre-committed to the data’s answer)
Personal scandal (sex, texts, personal graft) still ends American political careers; abuse-of- power scandal stopped carrying consequences. Phase 1 builds the dataset that confirms, refines, or refutes this. The phase-2 visualization presents whatever pattern the verified dataset actually shows — the thesis is a hypothesis being stress-tested, not an assumption baked into selection.
Scope
- Who: U.S. federal elected officials (President, VP, senators, representatives) and Senate-confirmed cabinet-level officers.
- When: scandal became public 1970-01-01 → today.
- What: credible public allegation of misconduct that drew national coverage (verifiable in ≥2 national outlets).
- Size: target 70–80 entries; both parties by construction.
Judges, U.S. Attorneys, White House staff (chiefs of staff, NSC/NSA staff, press secretaries, lawyers, campaign managers/treasurers), sub-cabinet appointees (deputy/assistant secretaries, agency chiefs below cabinet rank), state/local officials, and non-officeholder candidates are out of scope even where they appear on the source Wikipedia scandal lists — those lists are not themselves scoped to elected federal officials.
Seed method (anti-cherry-picking)
The candidate universe was pulled systematically — not by vibes or by memory — from four Wikipedia lists, fetched in full as raw wikitext to avoid summarization loss:
- List of federal political scandals in the United States
(
en.wikipedia.org/wiki/List_of_federal_political_scandals_in_the_United_States) — read in full from the Nixon administration section through the second Trump administration section (all 1970+ content), executive/legislative/judicial subsections. - List of American federal politicians convicted of crimes
(
en.wikipedia.org/wiki/List_of_American_federal_politicians_convicted_of_crimes) — full table extracted. - List of federal political sex scandals in the United States
(
en.wikipedia.org/wiki/List_of_federal_political_sex_scandals_in_the_United_States) — full chronological list, 1970+. - List of United States senators expelled or censured and List of United States representatives expelled, censured, or reprimanded — full disciplinary-action tables.
Every row pulled from these four sources was filtered by the scope criteria above and
independently verified before being written into the dataset. This produced a 187-row candidate
universe, of which 79 were selected for research and 78 were verified to the integrity bar below
(1 — thurmond-harassment — could not be verified and was dropped; see Exclusions). A limitation
of the method, not an oversight: because the universe draws only from these four list pages,
some well-known abuse-of-power/institutional-deception episodes documented elsewhere on
Wikipedia (e.g. Fast and Furious, Benghazi) do not appear in the candidate universe and were not
added from memory.
Inclusion criteria
Every selected row had to clear, independently:
- The officeholder scope test above (federal elected official or Senate-confirmed cabinet-level officer, at the time of the conduct).
- The 1970-01-01 date floor (scandal became public on or after this date).
- Plausible ≥2-national-outlet coverage — confirmed for real during dossier research, not assumed from the seed pull.
- The receipt integrity bar (below) — if the facts couldn’t be verified to this bar, the entry was dropped entirely (NOT-VERIFIABLE) rather than published with weakened sourcing.
Rubric (verbatim from the phase-1 design spec, fixed before research began)
Categories (primary; one of four)
- Personal conduct — sex, affairs, harassment, texts, personal behavior.
- Personal enrichment — bribery, graft, insider trading, misuse of funds for personal benefit.
- Abuse of power — using the office against opponents, elections, investigations, or institutions; includes obstruction.
- Institutional deception — lying to the public about war, policy, or coverups absent direct personal gain.
Each entry records a one-sentence categorization rationale. Genuinely mixed cases get one primary category plus an optional secondary tag (secondary must differ from primary).
Consequence Scale (ordinal; highest event within 5 years of the scandal breaking)
| Level | Meaning |
|---|---|
| 5 | Criminal conviction / prison |
| 4 | Resignation, removal, or withdrawal (from office, race, or nomination) |
| 3 | Formal sanction: indictment (regardless of verdict), House impeachment, censure, survived expulsion vote, loss of leadership/committee posts, court-ordered fine or civil judgment |
| 2 | Electoral defeat in the next cycle — recorded as correlation, explicitly not asserted as causation (methodology page states this) |
| 1 | No formal consequence; retained office or won the next election |
Take the maximum applicable level (e.g., resigned then convicted → 5). Each entry records the specific consequence event, its date, and a one-sentence rationale for the level.
Window clarification (amended 2026-07-17): the 5-year window bounds indirect consequences
(resignation, electoral defeat, informal fallout). Direct legal outcomes of the scandal’s own
conduct — indictment, conviction, or court judgment arising from that conduct — count
regardless of how long prosecution took. Rationale: the window exists to bound causal
attribution, and a conviction for the conduct is attributable by definition. Applied
symmetrically to all entries. (Applied concretely to trump-hushmoney: the 2018 revelation is
yearPublic, but the 2023 indictment and 2024 conviction count toward the level-5 score despite
the ~6-year lag, because they are direct legal outcomes of the underlying conduct, not indirect
fallout.)
Rubric interpretation notes (applied consistently across the dataset)
- The level-5 “conviction” OR-reading. The scale label reads “conviction/prison.” Read as
an OR, not an AND: a criminal conviction of any grade, including misdemeanor pleas without
incarceration, scores level 5. This includes court-accepted nolo contendere pleas (e.g.
agnew-bribery— no prison time, but a federal court’s acceptance of a nolo contendere plea produces a judgment of conviction) and misdemeanor guilty pleas with no incarceration (e.g.bowman-firealarm, a misdemeanor plea with a probation/eventual-dismissal provision). The scale’s two nouns are read as alternative triggers for the same level, not a conjunctive requirement that both be satisfied. - Level-2 correlation caveat. Every level-2 entry’s rationale explicitly states the consequence is recorded as correlation, not asserted causation — an electoral defeat following a scandal is not proof the scandal caused the defeat, and the dataset does not claim otherwise.
- Category corrections during research. Where the seed list’s proposed category didn’t hold
up under research, the dossier recategorized and documented why (e.g.
butz-remarkswas moved from institutional-deception to personal-conduct on review — the conduct was personal remarks by the official, not an institutional coverup).
Receipt integrity bar
Every receipt in every entry (2-3 per entry) had to clear, per the project’s written research rules (fixed before research began and applied to every entry):
- No fabrication. Never invent, guess, or approximate a URL, title, outlet, or date. If a fact couldn’t be verified, it was left out and the entry verdicted NOT-VERIFIABLE rather than stretched.
- Exact metadata. Exact published title (as it appears on the page), exact URL, correct outlet and date (YYYY-MM-DD).
- Live URL check. The URL must return 200/301/302 on a
curl -sIcheck with a realistic user agent. Non-bot-blocking outlets were preferred (NPR, Guardian, Politico, PBS, CBS, ABC, The Hill, AP via member stations); court records, DOJ press releases, and official congressional records (censure/expulsion resolutions, ethics reports) count as excellent receipts. - Concrete archive snapshot.
archiveUrlmust be a specific 14-digit-timestampweb.archive.orgsnapshot — never a*calendar URL — verified via the Wayback CDX API or created live if none existed. - Archive content verified. The archived snapshot was fetched and grepped for the exact cited title, not just checked for a 200 status; a snapshot that 200s but shows a 404 page or a search UI fails this check. Where Wayback replay was blocked from this network, the CDX status-200 timestamp plus a live-page title grep served as the documented fallback, logged as such in the dossier’s evidence log.
- Historical entries may use a high-quality retrospective from a major outlet as one of the 2-3 receipts; contemporaneous coverage is preferred where it exists.
Every dossier’s ## Evidence log records, per receipt, live status, which archive-verification
method was used, the title-grep match, and which specific receipt supports the dated consequence
event.
Exclusions
Two entries did not make the final 78-entry dataset despite being in-scope candidates at some point in the process:
petraeus-classified(David Petraeus, CIA Director, 2012) — dropped for scope: the CIA Director post was not cabinet-rank during Petraeus’s 2011–2012 tenure (that elevation began in 2017); the entry fails the Senate-confirmed cabinet-level test applied to every other cabinet-level entry in the dataset.thurmond-harassment(Strom Thurmond, U.S. Senator, proposed 1996) — dropped as NOT-VERIFIABLE: the seed row conflated two unrelated facts under one date. The 1996-public Murray/elevator allegation is real but thin — single-book-sourced, disputed, uninvestigated, and the alleged victim herself declined to characterize it as harassment when asked directly — and produced no formal consequence at any dated level 2-5. The second fact (a 1925 paternity case) is genuine but did not become public until December 2003, six months after Thurmond’s death, making it unusable as a consequence-bearing scandal entry for a serving officeholder. No dossier was forced past the integrity bar to fill the slot.
Seed-stage exclusions (108 candidates, not carried into research)
Beyond the two above, 108 of the 187 candidates identified in the systematic Wikipedia pull were never selected for research — the seed worksheet records a per-row rationale for every exclusion. They cluster into four categories:
- Out of scope (~16) — not a covered office (state/local, judicial branch, sub-cabinet appointee), predates the 1970-01-01 window, or the person was a candidate/nominee rather than a sitting officeholder at the time of the conduct.
- Duplicate-of a selected entry (~18) — same underlying scheme or event as an already- selected row (e.g. multiple Abscam or Koreagate participants beyond the anchor cases), kept once to avoid inflating a single episode into many dataset rows.
- Trimmed for balance (~28) — in-scope and individually defensible, but cut to hit the ~75-80 entry target and avoid over-weighting the 2005-2019 congressional-corruption wave. A reasonable future revision could swap these back in.
- Thin coverage (~34) — the ≥2-national-outlet bar was judged unlikely to clear on inspection (low-profile figures, single-state coverage, or disputed/unverifiable specifics), so these were not sent to research at all.
A note on category balance (not an exclusion, a property of the seed universe)
The verified 78-entry set skews toward personal-conduct and personal-enrichment relative to abuse-of-power and institutional-deception. This is not an artifact of the selection step — the four Wikipedia source lists are built around individual officeholder misconduct (bribery, affairs, ethics violations), which structurally produces many personal-conduct/personal- enrichment rows. Abuse-of-power and institutional-deception episodes (Watergate, Iran-Contra, Iraq WMD, the Plame affair, the U.S. Attorneys firings, Trump-Ukraine, Trump-Jan6, the Trump documents case, the Mayorkas impeachment) are structurally rarer and tend to implicate one scandal-defining principal (president, VP, AG) plus a long tail of aides who are out of scope under the officeholder-only filter — so each such episode contributes only 1-3 in-scope rows, versus the many discrete individuals a bribery ring like Abscam or Koreagate contributes. This imbalance is treated as a live input to the phase-1 thesis test, not smoothed over.
Dataset location
78 verified entries make up the site’s dataset, one schema-validated record per scandal — every field shown on the site (category, level, dates, receipts) comes from these records. Full research evidence (live-status checks, archive verification, title-grep matches, verdicts) is kept in a per-entry research dossier maintained alongside the dataset.
Congressional-discipline tier note: formal chamber discipline is scored at level 3 regardless of mechanism — full-chamber censure, denouncement, or reprimand votes, and Ethics Committee reprimands formally entered in the record (e.g. the Keating Five reprimand of Sen. Cranston) are treated as the same tier. The distinction between mechanisms is preserved in each entry’s consequence event text.
P1b expansion
Eleven entries — reagan-iran-contra, lance-omb-banking,
holder-fast-and-furious, ross-census-citizenship, burford-epa-contempt,
rice-benghazi, clinton-marc-rich-pardon, clinton-benghazi,
clinton-emails, lynch-tarmac, biden-documents — were added in a
targeted second pass, bringing the dataset to 89 entries.
Motivation
The phase-1 note on category balance (above) flagged abuse-of-power and
institutional-deception as structurally thin relative to personal-conduct and
personal-enrichment, an artifact of the four Wikipedia seed lists being built
around individual officeholder misconduct rather than institutional episodes.
P1b targeted that gap directly rather than waiting for another systematic
list pull: 9 of the 11 entries carry a primary category of abuse-of-power
(4) or institutional-deception (5); the remaining two (lance-omb-banking,
clinton-marc-rich-pardon) are personal-enrichment entries added for
era and party balance. The batch also corrected a
partisan-composition skew in the original 78-entry abuse-of-power/
institutional-deception cells (which leaned Republican — Watergate, Iran-
Contra, the Plame affair, Iraq WMD, the U.S. Attorneys firings) by adding
Democratic-administration entries in the same categories (Fast and Furious,
the Census citizenship question fight, the tarmac meeting, the Clinton
Benghazi and email entries, the Biden documents matter). Finally, the batch
closes four named blind spots the seed method’s own limitations note
(above) flagged as absent purely because they don’t appear on the four
source Wikipedia lists: Iran-Contra (Reagan and, via the existing
bush41-loop entry, Bush), Fast and Furious, the Census citizenship
question, and Benghazi (both the Rice talking-points angle and Clinton’s
State Department accountability).
Controller rubric decisions (binding, applied to this batch)
- Any-office level-2 rule, bush41-loop precedent. “Electoral defeat in
the next cycle” (level 2) applies to defeat in any federal election in
the cycle following the scandal, not only a defeat for the same office the
scandal implicated. This was first applied to
bush41-loop(Bush lost the presidency in 1992 over conduct from his vice presidency) and is applied here toclinton-emails(Clinton’s Nov. 2016 presidential loss, scored at level 2 with the mandatory correlation-not-causation qualifier, sourced to a 2017 NPR pollster review finding “at best mixed evidence” the Comey letter swung the race). - Consideration-vs-nomination boundary. Withdrawal from pre-nomination
consideration — never formally nominated — is not a level-4
(“resignation/removal/withdrawal”) event; only withdrawal from an
announced nomination counts.
gaetz-investigation(announced nominee for Attorney General, withdrew) scores level 4 on withdrawal;rice-benghazi(publicly floated as a possible Secretary of State pick, never formally nominated, withdrew from consideration Dec. 13, 2012) does not — the withdrawal is narrated in the summary but excluded from the level determination, leaving Rice at level 1 (no sanction on her actual office, later promoted to National Security Advisor). - Records-retention matched pair. Unauthorized post-office retention of
classified/government records is scored as abuse-of-power regardless of
whether obstruction is also present; the presence or absence of
obstruction is carried in the summary and consequence rationale, not the
category assignment. This governs the matched pair
trump-documents(obstruction found) /biden-documents(Special Counsel Hur found cooperation, not obstruction) — both score abuse-of-power, with the obstruction distinction stated in prose rather than expressed as a category difference. - Single-attribution rule for electoral defeats. An electoral defeat
attributes to at most one scandal entry — the most proximate/most-salient
scandal of that cycle — to prevent double-counting one election result
across multiple dataset rows. The Nov. 2016 loss attributes to
clinton-emailsonly (the Comey letter’s Oct. 28, 2016 revival of the story sits eleven days before the election, the most proximate causal candidate).clinton-benghaziremains level 1 despite falling inside the same 5-year window, and its summary carries a one-sentence cross-reference noting the 2016 loss is scored underclinton-emailsinstead.
OUT rejections
Three additional candidates considered for this batch were not added:
- Henry Kissinger (wiretapping of journalists and NSC staff) — rejected for office timing, the Petraeus precedent. The scandal became public via Seymour Hersh’s New York Times story “President Linked to Taps on Aides” (May 16, 1973); at that moment Kissinger held only the National Security Advisor role (White House staff, not Senate-confirmed). He was not confirmed as Secretary of State until Sept. 21, 1973 — right eventual title, wrong role at the moment of public disclosure.
- Mike Pence (classified documents found at his Indiana home, public Jan. 24, 2023; DOJ closed the investigation without charges June 2, 2023) — rejected for scope. At no point during the scandal’s public life was Pence a sitting officeholder or a declared candidate (he declared his 2024 campaign June 7, 2023, five days after the case closed) — the same former-officeholder exclusion class applied elsewhere in the dataset.
- Travelgate / Filegate (1993 White House travel-office firings; 1996 improper access to FBI background files) — rejected for person attachment. Independent Counsel Ray’s findings identify Hillary Clinton — First Lady, not an officeholder — as the “motivating force” in the travel-office firings, and place Filegate responsibility with sub-cabinet staff (the White House personnel-security office); neither episode attaches to Bill Clinton, the in-scope officeholder, so neither meets the rubric’s qualifying-person requirement.
Dataset size
The dataset now holds 89 verified entries. The expansion entries’ full research evidence is kept in the same per-entry dossier format as the original 78.