Insight · Method paper

Why most
impact numbers
do not survive
an audit

Four failure modes account for almost all of it. None requires bad faith — each follows naturally from designing measurement after delivery rather than before it.

Measurement Method 9 min read

A road cut through mountain terrain

Impact reporting has a credibility problem, and it is largely self-inflicted. The encouraging part is that the failures are systematic rather than random — which means they are preventable at design, and diagnosable in somebody else's report in about ten minutes.

Here is a test you can run on any impact report on your desk. Take the headline number and ask four questions. In our experience most published figures fail at least two.

Failure one: no baseline

The most common and the most fatal. A report states that 78% of participants now do something desirable. The question that decides whether this means anything is: how many did it before?

Without a starting position there is no change to report — only a state to describe. And a described state is compatible with the programme having achieved a great deal, nothing at all, or actively made things worse. The number is not wrong; it is simply uninterpretable.

The reason this happens is almost never dishonesty. It is sequencing. Measurement gets commissioned when the report is due, and by then the pre-intervention population no longer exists to be measured. A baseline is the one thing in a programme that genuinely cannot be added later.

Failure two: attrition absorbed into the result

A programme enrols 1,000 people. At endline, 600 are surveyed and 70% show the desired outcome. The report says 70%.

But 400 people left, and the reason they left is almost certainly correlated with the outcome. People for whom the intervention worked stay reachable. People who moved, disengaged, or for whom it did not work drop out disproportionately. Reporting 70% of the survivors while presenting it as 70% of participants is a selection effect presented as a finding.

The honest version reports both numbers and lets the reader see the gap.

420 of 1,000 is a materially different claim from 70%, and it is the claim the data supports. Attrition should appear as a count and a proportion, next to the result, every time.

Failure three: attribution claimed without a comparison

Something improved during the programme. The report says the programme improved it. Between those two statements sits every other thing that happened in that district over the same period — a good harvest, a government campaign, another organisation working two villages over, a change in the weather.

Where a control or comparison group is feasible, use one. Where it is not — and at district scale, on realistic timelines and budgets, it frequently is not — the correct move is to report the measured change and state plainly that the contribution cannot be isolated.

That reads as weaker. It is actually the stronger position, because it is the one that does not collapse when somebody asks the obvious question.

Failure four: reach reported in place of outcome

"We reached 40,000 people." Reached how? Present at an event is not reached. Handed a product is not reached. Living in a district where a campaign ran is definitely not reached.

Reach is an activity measure wearing the costume of an outcome measure. It is popular because it is cheap to produce and reliably large. It survives no scrutiny at all, because the first follow-up question — what changed for those 40,000 people — has no answer in the data.

The four questions, in order

01What was the measured starting position, and when was it taken?
02How many enrolled, how many were measured at endline, and where did the difference go?
03Against what comparison is the change being attributed?
04Is this an outcome, or an activity described as one?

Why this is becoming a commercial problem, not just an ethical one

For most of the sector's history, weak impact numbers carried no consequence. Narrative reporting was the norm and nobody audited a photograph.

That is changing, and the direction is visible in the disclosure regimes now forming around corporate sustainability reporting. In Saudi Arabia, the Capital Market Authority moved from voluntary ESG guidelines in 2019 to a structured Saudi Exchange framework in 2021 to binding requirements for green and sustainability-linked debt issuers in 2025, with ISSB-aligned standards in development and no confirmed mandatory date.12 In 2024, 94 listed companies published sustainability reports voluntarily.1

The relevance for anyone delivering programmes on behalf of those companies is direct. A figure that will eventually pass through an assurance process has requirements a narrative figure does not: a stated method, a defined boundary, traceable source records, and disclosed uncertainty. And a figure that has already been published without those things becomes a restatement risk rather than an achievement.

What a defensible result carries

Concretely, this is the set we assemble at the close of every programme — and the set we ask to see before we will cite anybody else's figure.

  • The problem statement agreed at scoping, including an explicit list of what was not going to be measurable. Written before delivery, so it cannot be revised to fit the result.
  • The indicator set, each mapped to a named external target rather than an internally-invented metric.
  • Baseline and endline: population, instrument, sample, date, method — for both, on the same terms.
  • Attrition and exclusions, with the reason for each exclusion.
  • Every adjustment and the reasoning behind it. Adjustments are legitimate; undisclosed adjustments are not.
  • Cost per outcome, modelled at design and restated against actuals at close. The gap between the two is one of the most informative numbers a programme produces, and almost nobody publishes it.
  • A programme review including what underperformed. A body of work with no failures in it has not been examined.

The test we hold ourselves to is deliberately awkward: the pack should be complete enough that a sceptical reader can disagree with our conclusion using our own data. If they cannot, we have published a claim rather than a finding.

The uncomfortable implication

Applying these four checks reliably produces smaller numbers. A programme that claimed 40,000 reached might report 3,200 measured, with 18% attrition, a stated change against baseline, and no attribution claim.

That is a less impressive paragraph and a far more valuable one — because it can be built on, compared, audited, and scaled with some confidence about what will happen. The larger number cannot be used for any of those things. It can only be repeated.

Sources

  1. Spectreco. Saudi Arabia ISSB sustainability reporting: Tadawul and CMA — regulatory sequencing and 2024 disclosure rates.
  2. Grant Thornton Saudi Arabia. ESG reporting in Saudi Arabia: preparing for IFRS S1 & S2 adoption.
  3. Saudi Exchange (Tadawul). ESG Disclosure Guidelines.

Our evidence standard

Four rules, applied to every programme, published whether the result was good or bad.

Read the standard The full method