Consequences Dataset — Methodology

Phase 1 of the “Consequences” section: a dataset of U.S. federal political scandals since 1970, each scored on a documented consequence scale with cited receipts. This page documents how the dataset was built, the rubric applied to every entry, and every exclusion.

Working thesis (pre-committed to the data’s answer)

Personal scandal (sex, texts, personal graft) still ends American political careers; abuse-of- power scandal stopped carrying consequences. Phase 1 builds the dataset that confirms, refines, or refutes this. The phase-2 visualization presents whatever pattern the verified dataset actually shows — the thesis is a hypothesis being stress-tested, not an assumption baked into selection.

Scope

Judges, U.S. Attorneys, White House staff (chiefs of staff, NSC/NSA staff, press secretaries, lawyers, campaign managers/treasurers), sub-cabinet appointees (deputy/assistant secretaries, agency chiefs below cabinet rank), state/local officials, and non-officeholder candidates are out of scope even where they appear on the source Wikipedia scandal lists — those lists are not themselves scoped to elected federal officials.

Seed method (anti-cherry-picking)

The candidate universe was pulled systematically — not by vibes or by memory — from four Wikipedia lists, fetched in full as raw wikitext to avoid summarization loss:

  1. List of federal political scandals in the United States (en.wikipedia.org/wiki/List_of_federal_political_scandals_in_the_United_States) — read in full from the Nixon administration section through the second Trump administration section (all 1970+ content), executive/legislative/judicial subsections.
  2. List of American federal politicians convicted of crimes (en.wikipedia.org/wiki/List_of_American_federal_politicians_convicted_of_crimes) — full table extracted.
  3. List of federal political sex scandals in the United States (en.wikipedia.org/wiki/List_of_federal_political_sex_scandals_in_the_United_States) — full chronological list, 1970+.
  4. List of United States senators expelled or censured and List of United States representatives expelled, censured, or reprimanded — full disciplinary-action tables.

Every row pulled from these four sources was filtered by the scope criteria above and independently verified before being written into the dataset. This produced a 187-row candidate universe, of which 79 were selected for research and 78 were verified to the integrity bar below (1 — thurmond-harassment — could not be verified and was dropped; see Exclusions). A limitation of the method, not an oversight: because the universe draws only from these four list pages, some well-known abuse-of-power/institutional-deception episodes documented elsewhere on Wikipedia (e.g. Fast and Furious, Benghazi) do not appear in the candidate universe and were not added from memory.

Inclusion criteria

Every selected row had to clear, independently:

  1. The officeholder scope test above (federal elected official or Senate-confirmed cabinet-level officer, at the time of the conduct).
  2. The 1970-01-01 date floor (scandal became public on or after this date).
  3. Plausible ≥2-national-outlet coverage — confirmed for real during dossier research, not assumed from the seed pull.
  4. The receipt integrity bar (below) — if the facts couldn’t be verified to this bar, the entry was dropped entirely (NOT-VERIFIABLE) rather than published with weakened sourcing.

Rubric (verbatim from the phase-1 design spec, fixed before research began)

Categories (primary; one of four)

  1. Personal conduct — sex, affairs, harassment, texts, personal behavior.
  2. Personal enrichment — bribery, graft, insider trading, misuse of funds for personal benefit.
  3. Abuse of power — using the office against opponents, elections, investigations, or institutions; includes obstruction.
  4. Institutional deception — lying to the public about war, policy, or coverups absent direct personal gain.

Each entry records a one-sentence categorization rationale. Genuinely mixed cases get one primary category plus an optional secondary tag (secondary must differ from primary).

Consequence Scale (ordinal; highest event within 5 years of the scandal breaking)

LevelMeaning
5Criminal conviction / prison
4Resignation, removal, or withdrawal (from office, race, or nomination)
3Formal sanction: indictment (regardless of verdict), House impeachment, censure, survived expulsion vote, loss of leadership/committee posts, court-ordered fine or civil judgment
2Electoral defeat in the next cycle — recorded as correlation, explicitly not asserted as causation (methodology page states this)
1No formal consequence; retained office or won the next election

Take the maximum applicable level (e.g., resigned then convicted → 5). Each entry records the specific consequence event, its date, and a one-sentence rationale for the level.

Window clarification (amended 2026-07-17): the 5-year window bounds indirect consequences (resignation, electoral defeat, informal fallout). Direct legal outcomes of the scandal’s own conduct — indictment, conviction, or court judgment arising from that conduct — count regardless of how long prosecution took. Rationale: the window exists to bound causal attribution, and a conviction for the conduct is attributable by definition. Applied symmetrically to all entries. (Applied concretely to trump-hushmoney: the 2018 revelation is yearPublic, but the 2023 indictment and 2024 conviction count toward the level-5 score despite the ~6-year lag, because they are direct legal outcomes of the underlying conduct, not indirect fallout.)

Rubric interpretation notes (applied consistently across the dataset)

Receipt integrity bar

Every receipt in every entry (2-3 per entry) had to clear, per the project’s written research rules (fixed before research began and applied to every entry):

  1. No fabrication. Never invent, guess, or approximate a URL, title, outlet, or date. If a fact couldn’t be verified, it was left out and the entry verdicted NOT-VERIFIABLE rather than stretched.
  2. Exact metadata. Exact published title (as it appears on the page), exact URL, correct outlet and date (YYYY-MM-DD).
  3. Live URL check. The URL must return 200/301/302 on a curl -sI check with a realistic user agent. Non-bot-blocking outlets were preferred (NPR, Guardian, Politico, PBS, CBS, ABC, The Hill, AP via member stations); court records, DOJ press releases, and official congressional records (censure/expulsion resolutions, ethics reports) count as excellent receipts.
  4. Concrete archive snapshot. archiveUrl must be a specific 14-digit-timestamp web.archive.org snapshot — never a * calendar URL — verified via the Wayback CDX API or created live if none existed.
  5. Archive content verified. The archived snapshot was fetched and grepped for the exact cited title, not just checked for a 200 status; a snapshot that 200s but shows a 404 page or a search UI fails this check. Where Wayback replay was blocked from this network, the CDX status-200 timestamp plus a live-page title grep served as the documented fallback, logged as such in the dossier’s evidence log.
  6. Historical entries may use a high-quality retrospective from a major outlet as one of the 2-3 receipts; contemporaneous coverage is preferred where it exists.

Every dossier’s ## Evidence log records, per receipt, live status, which archive-verification method was used, the title-grep match, and which specific receipt supports the dated consequence event.

Exclusions

Two entries did not make the final 78-entry dataset despite being in-scope candidates at some point in the process:

Seed-stage exclusions (108 candidates, not carried into research)

Beyond the two above, 108 of the 187 candidates identified in the systematic Wikipedia pull were never selected for research — the seed worksheet records a per-row rationale for every exclusion. They cluster into four categories:

A note on category balance (not an exclusion, a property of the seed universe)

The verified 78-entry set skews toward personal-conduct and personal-enrichment relative to abuse-of-power and institutional-deception. This is not an artifact of the selection step — the four Wikipedia source lists are built around individual officeholder misconduct (bribery, affairs, ethics violations), which structurally produces many personal-conduct/personal- enrichment rows. Abuse-of-power and institutional-deception episodes (Watergate, Iran-Contra, Iraq WMD, the Plame affair, the U.S. Attorneys firings, Trump-Ukraine, Trump-Jan6, the Trump documents case, the Mayorkas impeachment) are structurally rarer and tend to implicate one scandal-defining principal (president, VP, AG) plus a long tail of aides who are out of scope under the officeholder-only filter — so each such episode contributes only 1-3 in-scope rows, versus the many discrete individuals a bribery ring like Abscam or Koreagate contributes. This imbalance is treated as a live input to the phase-1 thesis test, not smoothed over.

Dataset location

78 verified entries make up the site’s dataset, one schema-validated record per scandal — every field shown on the site (category, level, dates, receipts) comes from these records. Full research evidence (live-status checks, archive verification, title-grep matches, verdicts) is kept in a per-entry research dossier maintained alongside the dataset.

Congressional-discipline tier note: formal chamber discipline is scored at level 3 regardless of mechanism — full-chamber censure, denouncement, or reprimand votes, and Ethics Committee reprimands formally entered in the record (e.g. the Keating Five reprimand of Sen. Cranston) are treated as the same tier. The distinction between mechanisms is preserved in each entry’s consequence event text.

P1b expansion

Eleven entries — reagan-iran-contra, lance-omb-banking, holder-fast-and-furious, ross-census-citizenship, burford-epa-contempt, rice-benghazi, clinton-marc-rich-pardon, clinton-benghazi, clinton-emails, lynch-tarmac, biden-documents — were added in a targeted second pass, bringing the dataset to 89 entries.

Motivation

The phase-1 note on category balance (above) flagged abuse-of-power and institutional-deception as structurally thin relative to personal-conduct and personal-enrichment, an artifact of the four Wikipedia seed lists being built around individual officeholder misconduct rather than institutional episodes. P1b targeted that gap directly rather than waiting for another systematic list pull: 9 of the 11 entries carry a primary category of abuse-of-power (4) or institutional-deception (5); the remaining two (lance-omb-banking, clinton-marc-rich-pardon) are personal-enrichment entries added for era and party balance. The batch also corrected a partisan-composition skew in the original 78-entry abuse-of-power/ institutional-deception cells (which leaned Republican — Watergate, Iran- Contra, the Plame affair, Iraq WMD, the U.S. Attorneys firings) by adding Democratic-administration entries in the same categories (Fast and Furious, the Census citizenship question fight, the tarmac meeting, the Clinton Benghazi and email entries, the Biden documents matter). Finally, the batch closes four named blind spots the seed method’s own limitations note (above) flagged as absent purely because they don’t appear on the four source Wikipedia lists: Iran-Contra (Reagan and, via the existing bush41-loop entry, Bush), Fast and Furious, the Census citizenship question, and Benghazi (both the Rice talking-points angle and Clinton’s State Department accountability).

Controller rubric decisions (binding, applied to this batch)

  1. Any-office level-2 rule, bush41-loop precedent. “Electoral defeat in the next cycle” (level 2) applies to defeat in any federal election in the cycle following the scandal, not only a defeat for the same office the scandal implicated. This was first applied to bush41-loop (Bush lost the presidency in 1992 over conduct from his vice presidency) and is applied here to clinton-emails (Clinton’s Nov. 2016 presidential loss, scored at level 2 with the mandatory correlation-not-causation qualifier, sourced to a 2017 NPR pollster review finding “at best mixed evidence” the Comey letter swung the race).
  2. Consideration-vs-nomination boundary. Withdrawal from pre-nomination consideration — never formally nominated — is not a level-4 (“resignation/removal/withdrawal”) event; only withdrawal from an announced nomination counts. gaetz-investigation (announced nominee for Attorney General, withdrew) scores level 4 on withdrawal; rice-benghazi (publicly floated as a possible Secretary of State pick, never formally nominated, withdrew from consideration Dec. 13, 2012) does not — the withdrawal is narrated in the summary but excluded from the level determination, leaving Rice at level 1 (no sanction on her actual office, later promoted to National Security Advisor).
  3. Records-retention matched pair. Unauthorized post-office retention of classified/government records is scored as abuse-of-power regardless of whether obstruction is also present; the presence or absence of obstruction is carried in the summary and consequence rationale, not the category assignment. This governs the matched pair trump-documents (obstruction found) / biden-documents (Special Counsel Hur found cooperation, not obstruction) — both score abuse-of-power, with the obstruction distinction stated in prose rather than expressed as a category difference.
  4. Single-attribution rule for electoral defeats. An electoral defeat attributes to at most one scandal entry — the most proximate/most-salient scandal of that cycle — to prevent double-counting one election result across multiple dataset rows. The Nov. 2016 loss attributes to clinton-emails only (the Comey letter’s Oct. 28, 2016 revival of the story sits eleven days before the election, the most proximate causal candidate). clinton-benghazi remains level 1 despite falling inside the same 5-year window, and its summary carries a one-sentence cross-reference noting the 2016 loss is scored under clinton-emails instead.

OUT rejections

Three additional candidates considered for this batch were not added:

Dataset size

The dataset now holds 89 verified entries. The expansion entries’ full research evidence is kept in the same per-entry dossier format as the original 78.

← consequences