Methodology

How Satyameter measures accuracy. This is a prototype: every verdict is AI-assessed and pending human review, and every verdict links to its evidence.

  1. The accuracy score

    Accuracy = accurate / (accurate + inaccurate + misleading), computed only over claims whose verdict carries a confidence of at least 0.7. Claims judged unverifiable are excluded from the denominator. The number of checked claims (N) is shown next to every score. No score is displayed when N is below 10 — those anchors sit below the line on the ledger with a dashed, needle-less meter.

  2. Why we rank by the confidence bound

    An anchor judged 9-of-10 accurate is not safely ahead of one judged 88-of-100 — the first has barely been measured. So we do not rank by raw accuracy. We compute a 95% Wilson confidence interval around each accuracy figure and rank by its lower bound: the score an anchor can defend even in the pessimistic case. A small sample yields a wide interval and a low bound, so anchors earn their rank by being checked more, not just by looking good on a handful of claims.

    The interval is drawn as a shaded band on each truth meter, and printed as CI 72–96% beside the checked count. As the verification campaign checks more claims, the band narrows and the ranking firms up.

    We also show scrutiny — checked claims as a fraction of all checkable ones (N checked of M checkable). A high accuracy over a low scrutiny is a provisional reading, not a verdict on the anchor.

  3. Reading the truth meter

    The meter is a 0–100 instrument scale. The needle sits at the anchor's accuracy; the shaded band around it is the 95% confidence interval, and the faint zones beneath the axis mark low (0–50), middling (50–80) and high (80–100) reading zones. It is a measurement, not a grade — follow any claim to its evidence and judge for yourself.

  4. Coverage profile

    The share of an anchor's extracted claims falling into each topic bucket — communal/religion, Pakistan/national-security, opposition-attack, governance/economy, crime, international, and others. For the leaderboard findings, topics are grouped into conflict & division (communal, security, opposition, crime, international) and development & governance (governance/economy, healthcare, education, infrastructure, agriculture, environment, science/tech). This is judgment-free counting, not part of the accuracy score.

  5. Who is held responsible

    For each claim we record the framing target — the party the anchor assigns responsibility to: individuals, systems & institutions, government, opposition, a minority community, a foreign actor, or the media. An anchor who blames the people involved frames differently from one who blames the system that enabled the issue. The dossier and topic pages direct-label the most frequent target.

  6. Evidence audit

    Every cited URL is fetched, archived, and status-recorded, so verdicts rest on reachable, snapshotted sources rather than unchecked citations. Each source shows its audit chip: ✓ reachable · archived, ✗ unreachable, or pending audit. When every source behind a verdict is dead, that verdict is excluded from scoring pending review.

  7. The pipeline

    1. Ingest — YouTube uploads on each network channel are scanned; episodes are matched to an anchor's show by title and filtered to a recent window.
    2. Transcribe — auto/uploaded captions are downloaded and cleaned into a transcript.
    3. Extract — an LLM pass extracts discrete claims made by the anchor, each with a topic tag, a factual/opinion/prediction class, and a checkability flag.
    4. Verify — each checkable factual claim is fact-checked against Indian fact-checkers (AltNews, BOOM Live, Factly, Newschecker, PIB Fact Check) and reputable primary sources; the verdict stores its evidence links.
    5. Enrich — each claim is tagged with a subtopic and a framing target for the coverage and responsibility analyses.
    6. Score — verdicts roll up into the per-anchor accuracy score and coverage profile.
  8. Honesty layer

    All verdicts are labelled AI-assessed — pending human review until a human reviews them. Evidence links are mandatory on every verdict. Scores and evidence are presented without characterization; readers can follow every link and judge for themselves.