Methodology
Every nstate scorecard follows this template. The methodology is the contract: if an analysis deviates from these standards, it is not an nstate scorecard.
seven sections, no exceptions
Every published scorecard must complete all seven sections. Sections left blank invalidate the scorecard.
Policy definition
Name, legislation or statutory basis, stated objectives (verbatim from official source), spending classification, temporal scope, and administering department. No paraphrase — exact official language only.
Observed data
Actual spending, headcount, or output. Every figure requires: primary source URL, data release date, metric definition, and the Python or R pipeline that fetched it. No figures from secondary reporting.
Counterfactual
What would have happened without the policy, or with an alternative policy. Method hierarchy (in order of preference): direct accounting, difference-in-differences, synthetic control, scenario modelling, descriptive only.
Impact dimensions
Fiscal impact, economic impact (efficiency/growth effects), delivery performance (targets met, quality, equity), and stated vs observed objectives. Each dimension scored separately.
Evidence quality
HIGH — randomised controlled trial, natural experiment, or peer-reviewed quasi-experimental evidence with pre-registration.
MEDIUM — credible observational study with appropriate controls, or official OBR/NAO model.
LOW — descriptive statistics without causal identification.
NOT MEASURABLE — insufficient data exists to score this dimension.
Agent benchmark (Phase 3)
For administrative processes only: estimated cost of current human delivery, estimated cost of deterministic software automation, estimated cost of AI-assisted automation, and evidence quality for each estimate. This section is blank until Phase 3.
Verdict
Three separate verdicts: fiscal, economic, and delivery. Each is: WASTE · MIXED · EFFECTIVE · INSUFFICIENT DATA. No single overall verdict — the dimensions are independent.
how we build alternatives
The counterfactual is the hardest part of any policy audit. We follow a strict hierarchy — the lowest level used must be stated explicitly.
Direct accounting
Where spending is directly fungible (e.g. a levy charged at a specific rate with a specific beneficiary), calculate the exact fiscal cost and compare to the next-best alternative using the same accounting standards.
Difference-in-differences
Compare outcomes in treated vs untreated units before and after the policy, controlling for pre-existing trends. Requires parallel trends assumption to be tested and reported.
Synthetic control
Construct a weighted combination of comparable countries or regions that match pre-treatment trends. Compare post-treatment divergence.
Scenario modelling
Where causal identification is not possible, use official forecasting models (OBR, OEF, OECD) with stated assumptions. Sensitivity ranges must be published.
Descriptive only
Where no counterfactual can be credibly constructed, describe the observed data only. Evidence quality is automatically LOW or NOT MEASURABLE. No verdict is issued.
what every scorecard must meet
no numbers without sources
Every figure requires a source URL, data release date, and the pipeline that fetched it. Figures from secondary media reporting are not permitted.
no editorialising in scored fields
Sections A–G use objective language. Opinion and advocacy belong in the "analysis" subsection only, clearly labelled as such.
corrections are mandatory and permanent
Errors are corrected promptly. The original version, the correction, and the reason are all published in the git history. Corrections are never silent.
right of reply
Any government department, programme manager, or academic may submit a counter-analysis via PR. If it meets methodology standards, it is published alongside the original.
no targeting individuals
Scorecards target programmes and policies, not individuals. Named officials appear only in a factual capacity (policy lead, minister responsible).
reproducible by default
Every scorecard has a corresponding data pipeline in the repository. A reader with a laptop must be able to reproduce every figure from primary sources.
what qualifies for a scorecard
Five requirements. All must be met before a scorecard enters the pipeline.
- A primary data source exists in an official statistical release (not press release, not media report).
- The policy has a stated objective against which delivery can be measured.
- A credible counterfactual exists at Level 3 or above.
- The fiscal scale exceeds £100m per year, or the policy is a priority area with unusually high public interest.
- The analysis team includes at least one independent reviewer with no prior stated position on this specific policy.
Want to submit a scorecard or challenge an existing one?
contribution guide →