# Honest Politics Methodology Technical Manual

## Status and relationship

This manual explains how to apply the Decision Standard. It contains technical methods and quality controls. It does not create additional policy stages, approval gates or mandatory records beyond the Decision Standard.

Where this manual offers several methods, the analyst must choose the method suited to the decision and record why.

## 1. Materiality and escalation



Scrutiny level should rise with:

- public cost;
- private cost;
- rights restriction;
- coercion;
- irreversibility;
- uncertainty;
- implementation complexity;
- geographic scale;
- number and vulnerability of people affected;
- constitutional effect;
- severe downside;
- dependence on untested technology or institutions.

The policy owner proposes the level. The methodology owner may require escalation.

Splitting a policy into smaller approvals to avoid scrutiny is prohibited.

## 1. Problem diagnosis

A problem definition should contain:

- baseline and trend;
- affected population;
- severity;
- causal map;
- institutional map;
- distribution;
- legal position;
- current policy;
- previous interventions;
- data limitations.

Useful tools include:

- causal diagrams;
- system maps;
- process mapping;
- administrative-data analysis;
- qualitative research;
- comparative analysis;
- failure analysis.

A causal map is a hypothesis, not proof.

## 1. Theory of change

The theory of change should describe:

- inputs;
- activities;
- outputs;
- behavioural or institutional mechanisms;
- intermediate outcomes;
- final outcomes;
- external conditions;
- possible harmful pathways.

Every important link should be classed as:

- established;
- plausible;
- disputed;
- unknown.

## 1. Option generation and shortlisting

The longlist should be broad enough to avoid solution capture.

Shortlisting may consider:

- legality;
- strategic fit;
- likely effectiveness;
- affordability;
- deliverability;
- rights;
- proportionality;
- risk;
- public acceptability;
- time;
- reversibility.

An option should not be removed solely because its political sponsor dislikes it.

Fatal flaws should be distinguished from remediable weaknesses.

## 1. Evidence by claim type

### Descriptive claims

Questions about scale, prevalence, distribution and trend.

Useful evidence:

- official statistics;
- administrative data;
- high-quality surveys;
- censuses;
- audits;
- transparent observational datasets.

### Causal claims

Questions about whether an action changes an outcome.

Useful evidence depends on context and may include:

- randomised trials;
- natural experiments;
- controlled comparisons;
- difference-in-differences;
- regression discontinuity;
- instrumental-variable designs;
- interrupted time series;
- longitudinal analysis;
- process and realist evaluation;
- triangulation.

Design quality and assumptions matter more than method labels.

### Mechanism claims

Questions about how and why effects arise.

Useful evidence:

- process evaluation;
- qualitative research;
- ethnography;
- interviews;
- behavioural data;
- implementation records;
- mixed methods.

### Forecast claims

Questions about future effects.

Useful evidence:

- validated models;
- reference classes;
- scenario analysis;
- market and behavioural evidence;
- expert elicitation;
- historical forecast performance.

### Delivery claims

Questions about whether institutions can implement.

Useful evidence:

- operational data;
- workforce and capacity analysis;
- procurement evidence;
- digital and service testing;
- frontline research;
- analogous delivery;
- pilot results.

### Value and legitimacy claims

Evidence can clarify views and consequences but cannot mechanically decide values.

Useful inputs:

- representative research;
- deliberative participation;
- constitutional analysis;
- public reasoning;
- elected judgement.

## 1. Evidence synthesis

The record should explain:

- search scope;
- inclusion and exclusion;
- source quality;
- relevance;
- consistency;
- heterogeneity;
- publication and selection risk;
- directness to the UK context;
- missing evidence;
- overall judgement.

A formal systematic review is not required for every decision. Major claims should nevertheless use a reproducible and challengeable search process.

## 1. Statistical interpretation

Report:

- effect size;
- uncertainty interval;
- base rate;
- sample and population;
- practical significance;
- heterogeneity;
- missing data;
- multiple testing where relevant;
- model assumptions.

Statistical significance is not a policy decision rule.

An inconclusive estimate is not proof of no effect.

Subgroup claims require appropriate power and a credible prior reason.

## 1. Forecasting and models

Material models require:

- named owner;
- purpose;
- inputs;
- assumptions;
- equations or logic;
- data sources;
- version;
- verification;
- validation;
- sensitivity;
- limitations;
- independent assurance proportionate to importance.

Use reference-class forecasting where comparable past delivery exists.

Adjust explicitly for optimism bias where relevant.

Separate:

- analytical uncertainty;
- implementation uncertainty;
- political uncertainty;
- exogenous scenario risk.

## 1. Cost and benefit appraisal

The appraisal should use real resource costs and avoid double counting.

Consider:

- public expenditure;
- revenue;
- transfer payments;
- private compliance;
- time;
- labour;
- capital;
- land;
- environmental effects;
- health and safety;
- distribution;
- productivity;
- wider system effects.

Fiscal impact and social value are not the same.

A tax receipt is not automatically a social benefit, and a benefit payment is not automatically a full social cost.

Use ranges and sensitivity rather than false precision.

## 1. Discounting and time

Apply current official appraisal guidance where government-standard analysis is intended.

The record should show:

- appraisal period;
- discount rate;
- reason;
- terminal or residual effects;
- long-term sensitivity;
- intergenerational and irreversible effects.

Discounting does not justify ignoring catastrophic or irreversible harm.

Current official discounting rules may change and should be checked when analysis is performed.

## 1. Distribution

Distributional analysis should examine relevant groups and places.

Do not select groups only after seeing results.

Report:

- absolute effect;
- relative effect;
- baseline condition;
- exposure to risk;
- ability to adapt;
- transition burden;
- access to remedies;
- cumulative effect with other policies.

Where relevant, assess protected characteristics and equality duties separately from broader distribution.

## 1. Behavioural and institutional response

At minimum test:

- target population;
- frontline workers;
- managers;
- suppliers;
- regulators;
- political actors;
- organised interests;
- fraudsters and hostile actors.

Ask:

- What becomes easier?
- What becomes more rewarding?
- What new avoidance route appears?
- What metric will people optimise?
- What information will decision-makers lose?
- What burden shifts elsewhere?
- Which organisation gains power or budget?
- What behaviour persists after the incentive ends?

## 1. Rights and proportionality

For a restriction of rights, record:

- legal basis;
- legitimate aim;
- evidence of need;
- affected right;
- severity and scope;
- less restrictive alternatives;
- safeguards;
- oversight;
- remedy;
- duration;
- review;
- discriminatory effect;
- hostile-successor risk.

Legal advice should be commissioned according to risk and should not be replaced by an internal checklist.

## 1. Public opinion and participation

Use the method suited to the question.

Representative polling estimates population opinion.

Open consultation discovers arguments and affected experiences.

Deliberative methods test considered judgement.

Expert submissions examine specialised claims.

Report each separately.

Opposition should trigger investigation, not rhetorical dismissal.

## 1. Independent challenge

Challenge should be:

- early enough to matter;
- matched to the claim;
- free to report disagreement;
- supported by access to evidence;
- subject to conflict disclosure.

Level 3 and Level 4 decisions normally require independent challenge.

Possible roles:

- analytical reviewer;
- subject expert;
- delivery reviewer;
- rights or legal reviewer;
- red team;
- affected-group reviewer.

A reviewer should not be described as independent where financial, personal or institutional dependence is material and unmanaged.

## 1. Delivery readiness

The delivery plan should identify:

- responsible organisation;
- senior responsible owner;
- frontline owner;
- powers;
- funding;
- workforce;
- technology;
- data;
- procurement;
- dependencies;
- transition;
- communications;
- complaints and redress;
- operational metrics;
- contingency;
- exit.

For major policies, an implementation-readiness review should occur before the final commitment or before irreversible rollout.

## 1. Pilots and test-and-learn

A pilot is justified where it can answer a material question.

State:

- uncertainty being tested;
- population and location;
- comparison;
- duration;
- sample;
- outcome;
- guardrails;
- decision threshold;
- scaling rule;
- stopping rule;
- learning publication.

A pilot should not be used where:

- harmful effects are irreversible;
- the sample cannot answer the question;
- the system will change before evaluation;
- the political decision to expand has already been made.

Test-and-learn may use:

- randomised rollouts;
- A/B tests;
- stepped-wedge designs;
- prototypes;
- sandboxes;
- matched comparisons;
- synthetic controls;
- rapid-cycle qualitative learning;
- adaptive trials;
- phased delivery.

The method should fit the risk and mechanism.

## 1. Evaluation

Evaluation should be designed during policy development.

It may include:

- process evaluation;
- impact evaluation;
- value-for-money evaluation;
- theory-based evaluation;
- systems and complexity approaches;
- distributional evaluation.

The evaluation question should determine the method.

Where complexity is high, evaluation should examine context, adaptation, interaction and mechanisms rather than searching only for one average treatment effect.

Learning should be used during delivery where safe and practical, not reserved for a final report.

## 1. Monitoring, review and exit

Monitoring should cover:

- output;
- outcome;
- distribution;
- cost;
- delivery;
- harm;
- public and frontline experience;
- fraud and misuse.

A policy should have:

- review date;
- decision owner;
- continuation rule;
- amendment rule;
- expansion rule;
- stop rule;
- sunset where appropriate;
- decommissioning plan where material.

Targets should include guardrails to reduce metric gaming.

## 1. International evidence and transferability

For a foreign example, record:

- problem and baseline;
- intervention;
- institutions;
- implementation;
- measured outcomes;
- causal confidence;
- costs and harms;
- economic context;
- population and geography;
- law and governance;
- culture and behaviour;
- delivery capacity;
- accompanying policies;
- time period;
- transferability judgement.

Classify the lesson as:

- directly informative;
- adaptable;
- mechanism only;
- cautionary;
- not transferable.

Avoid country-brand reasoning such as "Singapore proves" or "the Nordic model shows".

## 1. Deep uncertainty

Where probabilities or causal models are deeply uncertain:

- use multiple plausible futures;
- identify decisions robust across scenarios;
- preserve reversibility;
- stage commitments;
- build monitoring and triggers;
- avoid dependence on one central forecast;
- consider precaution where downside is severe and irreversible;
- consider the cost of delay.

## 1. Analytical quality assurance

Quality assurance should be proportionate and independent of the analyst who produced the work where risk justifies it.

Check:

- specification;
- data;
- code or calculations;
- logic;
- units;
- version;
- assumptions;
- sensitivity;
- outputs;
- documentation;
- reproducibility;
- security.

Material corrections should be logged.

## 1. Decision judgement

No metric automatically selects policy.

The final decision should distinguish:

- analytical conclusion;
- legal conclusion;
- delivery conclusion;
- public-legitimacy conclusion;
- political judgement.

Where the decision departs from the analytical recommendation, the reason should be recorded.

## 1. Public evidence card

The public record should state:

- position and status;
- problem;
- main options;
- preferred direction;
- principal evidence;
- benefits;
- costs;
- winners and losers;
- rights;
- uncertainty;
- strongest objection;
- delivery;
- success and review;
- what would change the position.

It should be readable without the technical annex.
