seven sections, no exceptions

Every published scorecard must complete all seven sections. Sections left blank invalidate the scorecard.

A

Policy definition

Name, legislation or statutory basis, stated objectives (verbatim from official source), spending classification, temporal scope, and administering department. No paraphrase — exact official language only.

B

Observed data

Actual spending, headcount, or output. Every figure requires: primary source URL, data release date, metric definition, and the Python or R pipeline that fetched it. No figures from secondary reporting.

C

Counterfactual

What would have happened without the policy, or with an alternative policy. Method hierarchy (in order of preference): direct accounting, difference-in-differences, synthetic control, scenario modelling, descriptive only.

D

Impact dimensions

Fiscal impact, economic impact (efficiency/growth effects), delivery performance (targets met, quality, equity), and stated vs observed objectives. Each dimension scored separately.

E

Evidence quality

HIGH — randomised controlled trial, natural experiment, or peer-reviewed quasi-experimental evidence with pre-registration.
MEDIUM — credible observational study with appropriate controls, or official OBR/NAO model.
LOW — descriptive statistics without causal identification.
NOT MEASURABLE — insufficient data exists to score this dimension.

F

Agent benchmark (Phase 3)

For administrative processes only: estimated cost of current human delivery, estimated cost of deterministic software automation, estimated cost of AI-assisted automation, and evidence quality for each estimate. This section is blank until Phase 3.

G

Verdict

Three separate verdicts: fiscal, economic, and delivery. Each is: WASTE · MIXED · EFFECTIVE · INSUFFICIENT DATA. No single overall verdict — the dimensions are independent.


how we build alternatives

The counterfactual is the hardest part of any policy audit. We follow a strict hierarchy — the lowest level used must be stated explicitly.

1 — strongest

Direct accounting

Where spending is directly fungible (e.g. a levy charged at a specific rate with a specific beneficiary), calculate the exact fiscal cost and compare to the next-best alternative using the same accounting standards.

2

Difference-in-differences

Compare outcomes in treated vs untreated units before and after the policy, controlling for pre-existing trends. Requires parallel trends assumption to be tested and reported.

3

Synthetic control

Construct a weighted combination of comparable countries or regions that match pre-treatment trends. Compare post-treatment divergence.

4

Scenario modelling

Where causal identification is not possible, use official forecasting models (OBR, OEF, OECD) with stated assumptions. Sensitivity ranges must be published.

5 — weakest

Descriptive only

Where no counterfactual can be credibly constructed, describe the observed data only. Evidence quality is automatically LOW or NOT MEASURABLE. No verdict is issued.


what every scorecard must meet

no numbers without sources

Every figure requires a source URL, data release date, and the pipeline that fetched it. Figures from secondary media reporting are not permitted.

no editorialising in scored fields

Sections A–G use objective language. Opinion and advocacy belong in the "analysis" subsection only, clearly labelled as such.

corrections are mandatory and permanent

Errors are corrected promptly. The original version, the correction, and the reason are all published in the git history. Corrections are never silent.

right of reply

Any government department, programme manager, or academic may submit a counter-analysis via PR. If it meets methodology standards, it is published alongside the original.

no targeting individuals

Scorecards target programmes and policies, not individuals. Named officials appear only in a factual capacity (policy lead, minister responsible).

reproducible by default

Every scorecard has a corresponding data pipeline in the repository. A reader with a laptop must be able to reproduce every figure from primary sources.


what qualifies for a scorecard

Five requirements. All must be met before a scorecard enters the pipeline.

  1. A primary data source exists in an official statistical release (not press release, not media report).
  2. The policy has a stated objective against which delivery can be measured.
  3. A credible counterfactual exists at Level 3 or above.
  4. The fiscal scale exceeds £100m per year, or the policy is a priority area with unusually high public interest.
  5. The analysis team includes at least one independent reviewer with no prior stated position on this specific policy.

Want to submit a scorecard or challenge an existing one?

contribution guide →