Over the past week, I taught the Impact Evaluation module for the MPhil programme in Evaluation at CREST. That meant Impact Evaluation was on my mind all week, and I got to dig into some analytical approaches behind different traditions.
What I tried to get across:
There is sometimes an assumption in impact evaluation that the most credible way to establish whether a programme had an effect is an experiment or quasi-experiment. These are powerful approaches. But they represent one way of building a causal argument. There are others.
A dice provides a surprisingly simple way of seeing the difference.
Let's start with a dice
Suppose I throw a fair six-sided dice five times.
Before I start, I can calculate the probability of getting exactly three sixes. It is about 3.2%.
Then I start throwing.
My first throw is a six.
Suddenly my chances of reaching exactly three sixes have changed. I now need two more sixes in the remaining four throws. The probability of that happening is about 11.6%.
Nothing about the dice has changed. I simply know something I didn't know before.
This is conditional probability: given what I now know, what is the probability of a particular outcome?
But let's use the same dice to think about three rather different ways of reasoning with evidence.
1. How surprising is what I observed?
I throw the dice five times and get three sixes.
Let's start by assuming that nothing unusual is going on. It is an ordinary fair dice.
Probability theory tells me what kinds of results I should expect if that assumption is true.
So I can ask:
If this really is a fair dice, how surprising is it that I got three sixes in five throws?
Quite surprising. Exactly three sixes would happen only about 3.2% of the time.
The reasoning goes something like this:
Assume nothing unusual is happening
↓
Work out what we would expect to observe under that assumption
↓
Compare what actually happened with that expectation
↓
Ask how surprising our observed result is
This is the intuition behind frequentist statistical testing.
Importantly, I am not asking, "What is the probability that my dice is fair?"
I am asking:
"If my dice is fair, how compatible is what I observed with that assumption?"
That distinction becomes important when we move into impact evaluation.
2. Given what I have seen, what should I now believe?
Now let's change the dice problem slightly.
This time, I don't know whether the dice is fair.
Before I throw it, I think it probably is. Most ordinary dice are.
Then I throw five times and get three sixes.
Hmm.
That doesn't prove that the dice is loaded. A fair dice can produce three sixes.
But perhaps I am now a little more suspicious than I was before.
So I throw another ten times. Suppose sixes continue appearing unusually often.
I become more suspicious.
I throw another twenty times and the pattern continues.
At each point I am asking:
Given what I believed before, and the evidence I have now observed, what should I believe now?
That is the basic intuition behind Bayesian reasoning:
What I believed before + new evidence → what I believe now.
And different pieces of evidence can change my confidence by different amounts.
Someone saying, "That dice looks funny", might shift my belief slightly.
Discovering that it came from a box labelled "Trick Dice" would shift it rather more.
Discovering that it has been deliberately weighted would shift it enormously.
So it isn't simply a question of accumulating lots of evidence.
We can also ask:
How much should this particular piece of evidence change my confidence in an explanation?
3. What combination produces the outcome?
Now let's change the game again.
Suppose there are several ways to win.
Perhaps:
two sixes + a five = WIN
but:
one six + three fives = WIN
also works.
And perhaps, after watching hundreds of games, we notice that nobody ever wins without at least one six.
Now I am asking a different kind of question.
Not:
"How probable is winning?"
but:
"What needs to come together for winning to happen?"
Perhaps there are several different combinations that are enough to produce a win.
Perhaps one condition is necessary but isn't enough on its own.
That is the basic intuition behind set-theoretic reasoning and approaches such as Qualitative Comparative Analysis (QCA).
So our dice has given us three questions:
Frequentist: If nothing unusual is happening, how surprising is what I observed?
Bayesian: Given what I believed before, how should the evidence I have observed change what I believe now?
Set-theoretic: What combination of conditions needs to come together to produce the outcome?
Now let's replace the dice with an education programme.
An education programme
Imagine we introduce a literacy programme into 100 schools.
The programme provides teacher training and ongoing coaching, with the intention that these improve classroom teaching and ultimately children's reading.
After two years, children in the programme schools have improved their reading scores by 10 points.
Excellent.
But have we established impact?
Not yet.
Children might have improved anyway. They are two years older. Teachers have gained experience. Perhaps government introduced new reading materials. Perhaps another education initiative was operating in the same districts.
So we need some way of asking:
What would probably have happened without our programme?
The experimental or quasi-experimental approach
Suppose we have a credible group of similar schools that did not receive the programme.
Children in those schools improved by 4 points.
Our programme schools improved by 10 points.
That gives us an estimated programme effect of 6 points.
This is where counterfactual reasoning gives experimental and quasi-experimental designs their power.
We have tried to construct a credible estimate of what would have happened to programme participants in the absence of the programme.
But there is still uncertainty.
Even if the programme had absolutely no effect, we would not expect two groups of children to produce precisely identical scores. There is ordinary variation in the world and in our samples.
And this is where we can return directly to our dice.
With the dice we asked:
If the dice really is fair, how surprising is it that I observed three sixes?
In the education evaluation we can ask:
If the programme really had no effect, how surprising would it be to observe a difference of 6 points or more between these groups?
The logic is the same:
Assume there is no programme effect
↓
Work out the kinds of differences we would expect under that assumption
↓
Compare our observed 6-point difference with that expectation
↓
Ask how surprising our result would be if there really were no effect
This is the frequentist statistical reasoning commonly used alongside experimental and quasi-experimental impact evaluations.
There is an important distinction here.
The counterfactual and the statistical test are not the same thing.
The experimental or quasi-experimental design helps us construct the causal comparison: what would probably have happened without the programme?
Frequentist statistical inference is one way of assessing the uncertainty around the effect we estimate from that comparison.
Indeed, an RCT could be analysed using Bayesian statistics instead.
But now suppose we cannot construct a credible counterfactual.
Does that mean we can say nothing rigorous about whether the programme contributed to the observed improvement?
I don't think so.
Contribution analysis and process tracing: How confident are we in the causal explanation?
Suppose the programme was deliberately introduced into the weakest districts.
Schools received different amounts of support. Government introduced a new reading policy at the same time. Other organisations were working in some of the same schools.
Finding a genuinely comparable group may be extremely difficult.
But we have a theory about how our programme was supposed to produce change:
Training
↓
Improved teacher knowledge
↓
Coaching helps teachers put that knowledge into practice
↓
Classroom teaching changes
↓
Children's reading improves
Now we can investigate that explanation.
Did teachers actually learn the intended practices?
Did their classroom practice subsequently change?
Were the changes the particular practices targeted by the programme?
What role did coaching play?
Did teachers receiving stronger coaching change more?
Did the aspects of children's reading that should have responded to the changed teaching actually improve?
Did the sequence of changes happen in the order our theory predicts?
And what else could explain what we observed?
Perhaps the government policy is the real explanation. Perhaps another literacy programme was responsible. Perhaps reading improved everywhere.
Each piece of evidence can strengthen or weaken our confidence in our explanation.
This takes us back to our possibly loaded dice.
We are asking:
Given everything we have now observed, how confident should we be in this particular explanation of what happened?
That kind of reasoning has a strong affinity with process tracing and, somewhat more loosely, contribution analysis.
We aren't pretending we have constructed a counterfactual when we haven't.
But neither are we simply observing that outcomes improved and claiming credit.
We are systematically testing the causal story and competing explanations.
QCA and realist approaches: Under what conditions does it work?
Now suppose something else happens.
Our average result hides enormous variation.
Some schools improve dramatically. Some improve a little. Others don't improve at all.
Perhaps the interesting question isn't simply:
"Did the programme work?"
Perhaps it is:
"What needs to come together for the programme to work?"
When we look across schools, we might discover one pathway to success:
Good coaching + strong school leadership + adequate materials → improved reading
But another group of successful schools might have:
Good coaching + strong teacher collaboration + district support → improved reading
There may be more than one route to the same outcome.
This is the territory of QCA and set-theoretic reasoning.
Realist evaluation approaches the problem somewhat differently. It asks us to investigate:
What works, for whom, in what circumstances, and why?
Both draw our attention to something that an average programme effect can easily obscure:
Context matters, and programmes may produce outcomes through different pathways in different circumstances.
So which one establishes impact?
This is where I think we sometimes get ourselves into trouble.
We talk about the "strongest" evaluation design as though all impact evaluations are trying to answer exactly the same causal question.
They aren't.
A good experimental or quasi-experimental evaluation can provide powerful evidence that:
People exposed to the programme experienced better outcomes than comparable people who were not exposed to it.
Contribution analysis or process tracing can investigate:
How confident are we that the programme contributed to the observed change through the causal processes we expected, rather than through plausible alternative explanations?
QCA can investigate:
Which combinations of programme and contextual conditions are associated with the outcome, and are there different pathways to success?
A realist evaluation can investigate:
What works, for whom, under what circumstances, and through what mechanisms?
These approaches do not make identical causal claims.
And that is the point.
If what we really need is a credible estimate of the difference an intervention made compared with what would otherwise have happened, then a good experimental or quasi-experimental design may be exactly what we need.
But sometimes that isn't the most important question.
Sometimes we need to understand whether and how a programme contributed to change in a complex environment.
Sometimes we need to know why it worked in 12 schools and not in eight others.
Sometimes we need to understand which combination of programme support and contextual conditions produces success.
And sometimes we need several of these answers.
So perhaps the first question in designing an impact evaluation shouldn't be:
"Can we do a quasi-experiment?"
It should be:
"What exactly do we need to know about whether, how, for whom and under what circumstances this programme made a difference?"
Then we can decide what kind of evidence — and what kind of causal reasoning — gives us the most credible answer.








