Sunday, September 13, 2026

Analytical approaches to establish causality

Over the past week, I taught the Impact Evaluation module for the MPhil programme in Evaluation at CREST. That meant Impact Evaluation was on my mind all week, and I got to dig into some analytical approaches behind different traditions. 

 What I tried to get across: 

There is sometimes an assumption in impact evaluation that the most credible way to establish whether a programme had an effect is an experiment or quasi-experiment. These are powerful approaches. But they represent one way of building a causal argument. There are others.

A dice provides a surprisingly simple way of seeing the difference.

Let's start with a dice

Suppose I throw a fair six-sided dice five times.

Before I start, I can calculate the probability of getting exactly three sixes. It is about 3.2%.

Then I start throwing.

My first throw is a six.

Suddenly my chances of reaching exactly three sixes have changed. I now need two more sixes in the remaining four throws. The probability of that happening is about 11.6%.

Nothing about the dice has changed. I simply know something I didn't know before.

This is conditional probability: given what I now know, what is the probability of a particular outcome?

But let's use the same dice to think about three rather different ways of reasoning with evidence.

1. How surprising is what I observed?


I throw the dice five times and get three sixes.

Let's start by assuming that nothing unusual is going on. It is an ordinary fair dice.

Probability theory tells me what kinds of results I should expect if that assumption is true.

So I can ask:

If this really is a fair dice, how surprising is it that I got three sixes in five throws?

Quite surprising. Exactly three sixes would happen only about 3.2% of the time.

The reasoning goes something like this:

Assume nothing unusual is happening

Work out what we would expect to observe under that assumption

Compare what actually happened with that expectation

Ask how surprising our observed result is

This is the intuition behind frequentist statistical testing.

Importantly, I am not asking, "What is the probability that my dice is fair?"

I am asking:

"If my dice is fair, how compatible is what I observed with that assumption?"

That distinction becomes important when we move into impact evaluation.

2. Given what I have seen, what should I now believe?


Now let's change the dice problem slightly.

This time, I don't know whether the dice is fair.

Before I throw it, I think it probably is. Most ordinary dice are.

Then I throw five times and get three sixes.

Hmm.

That doesn't prove that the dice is loaded. A fair dice can produce three sixes.

But perhaps I am now a little more suspicious than I was before.

So I throw another ten times. Suppose sixes continue appearing unusually often.

I become more suspicious.

I throw another twenty times and the pattern continues.

At each point I am asking:

Given what I believed before, and the evidence I have now observed, what should I believe now?

That is the basic intuition behind Bayesian reasoning:

What I believed before + new evidence → what I believe now.

And different pieces of evidence can change my confidence by different amounts.

Someone saying, "That dice looks funny", might shift my belief slightly.

Discovering that it came from a box labelled "Trick Dice" would shift it rather more.

Discovering that it has been deliberately weighted would shift it enormously.

So it isn't simply a question of accumulating lots of evidence.

We can also ask:

How much should this particular piece of evidence change my confidence in an explanation?

3. What combination produces the outcome?



Now let's change the game again.

Suppose there are several ways to win.

Perhaps:

two sixes + a five = WIN

but:

one six + three fives = WIN

also works.

And perhaps, after watching hundreds of games, we notice that nobody ever wins without at least one six.

Now I am asking a different kind of question.

Not:

"How probable is winning?"

but:

"What needs to come together for winning to happen?"

Perhaps there are several different combinations that are enough to produce a win.

Perhaps one condition is necessary but isn't enough on its own.

That is the basic intuition behind set-theoretic reasoning and approaches such as Qualitative Comparative Analysis (QCA).

So our dice has given us three questions:

Frequentist: If nothing unusual is happening, how surprising is what I observed?

Bayesian: Given what I believed before, how should the evidence I have observed change what I believe now?

Set-theoretic: What combination of conditions needs to come together to produce the outcome?

Now let's replace the dice with an education programme.

An education programme

Imagine we introduce a literacy programme into 100 schools.

The programme provides teacher training and ongoing coaching, with the intention that these improve classroom teaching and ultimately children's reading.

After two years, children in the programme schools have improved their reading scores by 10 points.



Excellent.

But have we established impact?

Not yet.

Children might have improved anyway. They are two years older. Teachers have gained experience. Perhaps government introduced new reading materials. Perhaps another education initiative was operating in the same districts.

So we need some way of asking:

What would probably have happened without our programme?

The experimental or quasi-experimental approach

Suppose we have a credible group of similar schools that did not receive the programme.

Children in those schools improved by 4 points.

Our programme schools improved by 10 points.

That gives us an estimated programme effect of 6 points.

This is where counterfactual reasoning gives experimental and quasi-experimental designs their power.

We have tried to construct a credible estimate of what would have happened to programme participants in the absence of the programme.

But there is still uncertainty.

Even if the programme had absolutely no effect, we would not expect two groups of children to produce precisely identical scores. There is ordinary variation in the world and in our samples.

And this is where we can return directly to our dice.

With the dice we asked:

If the dice really is fair, how surprising is it that I observed three sixes?

In the education evaluation we can ask:

If the programme really had no effect, how surprising would it be to observe a difference of 6 points or more between these groups?

The logic is the same:

Assume there is no programme effect

Work out the kinds of differences we would expect under that assumption

Compare our observed 6-point difference with that expectation

Ask how surprising our result would be if there really were no effect

This is the frequentist statistical reasoning commonly used alongside experimental and quasi-experimental impact evaluations.

There is an important distinction here.

The counterfactual and the statistical test are not the same thing.

The experimental or quasi-experimental design helps us construct the causal comparison: what would probably have happened without the programme?

Frequentist statistical inference is one way of assessing the uncertainty around the effect we estimate from that comparison.

Indeed, an RCT could be analysed using Bayesian statistics instead.

But now suppose we cannot construct a credible counterfactual.

Does that mean we can say nothing rigorous about whether the programme contributed to the observed improvement?

I don't think so.

Contribution analysis and process tracing: How confident are we in the causal explanation?

Suppose the programme was deliberately introduced into the weakest districts.

Schools received different amounts of support. Government introduced a new reading policy at the same time. Other organisations were working in some of the same schools.

Finding a genuinely comparable group may be extremely difficult.

But we have a theory about how our programme was supposed to produce change:

Training

Improved teacher knowledge

Coaching helps teachers put that knowledge into practice

Classroom teaching changes

Children's reading improves

Now we can investigate that explanation.

Did teachers actually learn the intended practices?

Did their classroom practice subsequently change?

Were the changes the particular practices targeted by the programme?

What role did coaching play?

Did teachers receiving stronger coaching change more?

Did the aspects of children's reading that should have responded to the changed teaching actually improve?

Did the sequence of changes happen in the order our theory predicts?

And what else could explain what we observed?

Perhaps the government policy is the real explanation. Perhaps another literacy programme was responsible. Perhaps reading improved everywhere.

Each piece of evidence can strengthen or weaken our confidence in our explanation.

This takes us back to our possibly loaded dice.

We are asking:

Given everything we have now observed, how confident should we be in this particular explanation of what happened?

That kind of reasoning has a strong affinity with process tracing and, somewhat more loosely, contribution analysis.

We aren't pretending we have constructed a counterfactual when we haven't.

But neither are we simply observing that outcomes improved and claiming credit.

We are systematically testing the causal story and competing explanations.

QCA and realist approaches: Under what conditions does it work?

Now suppose something else happens.

Our average result hides enormous variation.

Some schools improve dramatically. Some improve a little. Others don't improve at all.

Perhaps the interesting question isn't simply:

"Did the programme work?"

Perhaps it is:

"What needs to come together for the programme to work?"

When we look across schools, we might discover one pathway to success:

Good coaching + strong school leadership + adequate materials → improved reading

But another group of successful schools might have:

Good coaching + strong teacher collaboration + district support → improved reading

There may be more than one route to the same outcome.

This is the territory of QCA and set-theoretic reasoning.

Realist evaluation approaches the problem somewhat differently. It asks us to investigate:

What works, for whom, in what circumstances, and why?

Both draw our attention to something that an average programme effect can easily obscure:

Context matters, and programmes may produce outcomes through different pathways in different circumstances.

So which one establishes impact?

This is where I think we sometimes get ourselves into trouble.

We talk about the "strongest" evaluation design as though all impact evaluations are trying to answer exactly the same causal question.

They aren't.

A good experimental or quasi-experimental evaluation can provide powerful evidence that:

People exposed to the programme experienced better outcomes than comparable people who were not exposed to it.

Contribution analysis or process tracing can investigate:

How confident are we that the programme contributed to the observed change through the causal processes we expected, rather than through plausible alternative explanations?

QCA can investigate:

Which combinations of programme and contextual conditions are associated with the outcome, and are there different pathways to success?

A realist evaluation can investigate:

What works, for whom, under what circumstances, and through what mechanisms?

These approaches do not make identical causal claims.

And that is the point.

If what we really need is a credible estimate of the difference an intervention made compared with what would otherwise have happened, then a good experimental or quasi-experimental design may be exactly what we need.

But sometimes that isn't the most important question.

Sometimes we need to understand whether and how a programme contributed to change in a complex environment.

Sometimes we need to know why it worked in 12 schools and not in eight others.

Sometimes we need to understand which combination of programme support and contextual conditions produces success.

And sometimes we need several of these answers.

So perhaps the first question in designing an impact evaluation shouldn't be:

"Can we do a quasi-experiment?"

It should be:

"What exactly do we need to know about whether, how, for whom and under what circumstances this programme made a difference?"

Then we can decide what kind of evidence — and what kind of causal reasoning — gives us the most credible answer.


AI disclosure: I used ChatGPT as a thinking and writing aid while developing this post — to explore the distinctions between the different approaches, test and refine the examples, check my explanations, and help edit the final text. The argument, interpretation and final choices are my own.


AI made me worry about transitions

I am worried about AI. 'Cause researchers wrote a scary paper with scary words like "autonomous replication and adaptation".

I'm worried about increasingly capable systems becoming more autonomous than we intended. I'm worried about what bad actors might do with them. I'm worried about the concentration of extraordinary power in a handful of companies and countries. And I'm also worried about the less spectacular possibility that we are living through an enormous technology and investment bubble which bursts.

I don't know which of these risks is most likely. Nobody really does.

But thinking about AI led me to a slightly different question.



What do we do when we know a system is changing, but we cannot know what it is changing into?

Because AI isn't the only place where this feels relevant.

We are seeing profound changes in how philanthropy and development are funded. Governments are under fiscal pressure. Traditional development aid is shifting and drying up. Foundations are reconsidering their roles. Organisations are having to rethink sustainability. In South Africa, we are simultaneously worrying about institutional capacity, inequality, social cohesion and whether our democratic institutions are resilient enough to withstand future shocks.

Different transitions, certainly. But they pose a similar problem.

We can identify futures we would prefer. We can work towards them. But none of us controls the system well enough to produce them.

So where is our agency?

I have found myself increasingly attracted to a fairly simple answer:

Perhaps we can seed the transition with patterns that we would like the future system to have available.


2. What does it mean to seed a transition?

Consider a donor deciding to put substantial money behind an issue.

The obvious intervention is the money. It allows organisations to do things.

But something else happens too.

By deciding that this issue matters, the donor helps set an agenda. Other organisations pay attention. Researchers become interested. Government officials engage. Other funders may follow. Of course, agenda setting is an exercise of power. A funder can seed patterns we think are desirable, but it can just as easily reinforce patterns we don't. The point isn't that seeding is inherently good. It is that it happens — and perhaps we should be more deliberate about it.

Now suppose the donor does something more deliberate. Instead of funding a collection of separate projects, it creates incentives for collaboration. Organisations share information, work across their usual boundaries and collectively make sense of what they are learning.

And then suppose the donor introduces another practice: system-wide sensing.

The collaborators don't only monitor whether their projects are delivering. They continually look outward.

What is changing in the wider system? Where are new actors appearing? What are communities experiencing? Where is resistance emerging? Where are unexpected opportunities? What are other organisations learning? What are we not seeing?

The initiative itself may eventually end.

But imagine that some of those practices remain.

Another funder starts convening its partners in the same way. An NGO takes the sensing practice into another programme. A government department adopts a similar mechanism. People who met through the collaboration continue working together.

What has been replicated isn't necessarily the programme.

It is a pattern:

focus attention → bring different actors together → collaborate around a shared problem → sense what is happening across the system → learn and adapt.

That is what I mean by a seed.

And it made me wonder: could deliberately creating good patterns be a sensible strategy when we don't know what the future system will look like?

There is actually quite a lot of theory suggesting that it might be.


3. Finding agency without pretending we have control

The idea that first helped me was path dependence.

The basic proposition is intuitive: where a system can go next is partly shaped by the path it has already travelled.

A new system doesn't arrive on a blank page. It inherits institutions, relationships, capabilities, technologies, habits, inequalities, values and ways of working.

That means what exists before and during a transition matters.

If collaboration already exists, collaboration is available to the emerging system.

If trusted networks exist, they are available.

If institutions have developed ways of sensing and adapting, those capabilities are available.

If everything has become fragmented, competitive and distrustful, those patterns are available too.

We are seeding future systems whether we intend to or not.

The interesting question is whether we can become more deliberate about which patterns we strengthen.

THEORY BOX: Path dependence

Social systems have histories. Once particular institutions, technologies, rules or behaviours become established, they can reinforce themselves and make some future trajectories easier than others.

The future isn't predetermined. But it doesn't begin from scratch either.

Why it matters: If today's patterns influence tomorrow's possible pathways, then creating and strengthening desirable patterns now can matter even when we cannot predict the eventual outcome. 

        Pierson, P. (2000). “Increasing Returns, Path Dependence, and the Study of Politics.” American             Political Science Review, 94(2), 251–267. 


4. What kind of patterns do we want?


This is where I find the latest Indlulamithi South Africa Scenarios 2035 particularly useful.

They give us three possible South African futures through wonderfully evocative bird metaphors.

There is Hadeda Home: the Recrimination Nation. Noisy, anxious, defensive and fragmented.

There is Vulture Culture: the Desperation Nation. A darker trajectory of institutional deterioration, insecurity and capture.

And there is Weaver Work: the Cooperation Nation.

I love the weaver metaphor.

A sociable weaver doesn't become resilient because one exceptionally powerful bird builds the perfect nest. Many birds contribute to a shared structure. The collective structure, in turn, gives them greater resilience.

That raises an interesting question.

If we want a more cohesive and resilient society, what Weaver-like patterns should we be putting into our institutions now?

Collaboration across institutional boundaries.

Relationships that survive individual projects.

Institutions that can learn and adapt.

Ways of sharing information without pretending that there is only one interpretation of it.

Collective mechanisms for sensing what is happening across a system.

Participation that gives people genuine agency.

Accountability.

Ways of dealing with disagreement without retreating into our respective laagers.

Perhaps even the mundane practice of people who don't usually work together learning how to do so.

These aren't outcomes like social cohesion or resilience.

They are patterns from which cohesion and resilience might emerge.


5. Meadows: preserve the system's ability to evolve

This takes me to Donella Meadows.

One of the most powerful leverage points she identifies is a system's capacity to self-organise.

And self-organisation requires diversity and experimentation.

That suggests something slightly counter-intuitive for funders and evaluators.

In an uncertain transition, our job may not always be to identify the intervention that works, standardise it and scale it.

Sometimes we may need to maintain a sufficiently rich ecology of good possibilities.

Try things.

Protect experimentation.

Learn.

Allow different institutional arrangements to develop.

Keep alternatives alive.

In other words: seed.

THEORY BOX: Donella Meadows and self-organisation

Meadows argued that changing things like budgets and targets is relatively weak systems intervention. Much stronger leverage can come from changing information flows, rules, goals and the capacity of a system to self-organise.

Self-organisation depends partly on having diversity from which new structures and behaviours can emerge.

My takeaway: Don't only optimise the system we have. Preserve its ability to become something else.

        Meadows, D.H. (1999). Leverage Points: Places to Intervene in a System. Hartland, VT: The Sustainability Institute. 

          Abson, D.J. et al. (2017). “Leverage points for sustainability transformation.” Ambio, 46, 30–39. 


6. Transition theory: today's little experiments may be tomorrow's alternatives

Transition theory gives us another useful idea.

Established systems are hard to change because so many things hold them together: regulations, institutions, technologies, money, professional practices, skills and expectations.

But around their edges are smaller experiments — what transition scholars call niches.

Most remain small.

Some disappear.

But some survive long enough for people to learn from them, build networks around them and develop them into credible alternatives.

Then something disrupts the established system.

Suddenly, an alternative that previously looked marginal may become useful.

THEORY BOX: Transition theory

One influential approach distinguishes between:

Landscape: large external pressures and changes.

Regime: the established system of institutions, rules, technologies and practices.

Niches: spaces in which alternatives can develop.

When pressure destabilises the existing regime, sufficiently developed alternatives may suddenly have opportunities to spread.

Why it matters: A small experiment doesn't have to transform the current system to be worthwhile. It may be developing a pattern or capability that becomes important under different conditions.

        Geels, F.W. & Schot, J. (2007). “Typology of sociotechnical transition pathways.” Research Policy, 36(3), 399–417. An author manuscript is openly available through the University of Manchester repository 

Geels, F.W. (2026). “The Multi-level Perspective on Sustainability Transitions: Background, Overview and Current Research Topics.” In Introduction to Sustainability Transitions Research. Cambridge University Press. The whole volume is open access


7. Patton: transformation happens through connections

Michael Quinn Patton approaches this from an evaluation perspective through his idea of a Theory of Transformation.

A programme can reasonably have a Theory of Change:

If we do these things, through these mechanisms, we expect these outcomes.

But transformation doesn't work quite like that.

Nobody runs “the system”.

Transformation involves multiple actors, interventions, institutions and changes interacting across different levels.

And that brings Patton to an idea I find particularly useful: interconnectedness momentum.

Think back to our donor.

Perhaps the most significant thing it achieved wasn't its programme outcomes.

Perhaps it helped establish a new agenda.

Perhaps organisations that previously worked separately formed relationships.

Perhaps system-wide sensing became normal practice.

Perhaps those organisations started influencing other organisations.

Perhaps another funder adopted the approach.

Perhaps people carried those practices into completely different institutions.

The original intervention hasn't necessarily “scaled”.

The pattern has travelled.

THEORY BOX: Theory of Transformation

A Theory of Change asks how an intervention is expected to produce outcomes.

A Theory of Transformation asks how multiple changes across actors, organisations, sectors and levels might interact to contribute to a much larger transformation.

Nobody controls that transformation.

Patton therefore draws attention to interconnectedness momentum: whether growing connections and interactions between initiatives are creating momentum towards transformation.

        Patton, M.Q. (2021). “Evaluation Criteria for Evaluating Transformation: Implications for the Coronavirus Pandemic and the Global Climate Emergency.” American Journal of Evaluation, 42(1), 53–89 

That suggests a different set of evaluation questions.

Not only:

Did the intervention work?

But also:

What patterns did it introduce?

Which survived?

Who picked them up?

How did they change as they travelled?

What did they connect with?

Are those connections becoming self-sustaining?

Is something happening in the wider system that wasn't happening before?


8. This has to extend beyond evaluation — and beyond development

I started thinking about this because I was frightened about AI.

Then I recognised it in philanthropy.

But neither evaluation nor philanthropy is big enough.

If we really are living through a period of multiple overlapping transitions, patterns are being created everywhere: in schools, universities, families, businesses, government departments, communities, professional associations, civil society, technology companies and the media.

Some will reinforce Hadeda patterns.

Some may create Vulture patterns.

And some may be Weaver patterns.

That gives us a different way of thinking about agency.

We don't all have to be working on the same intervention.

We don't even have to agree on exactly what the future will look like.

But perhaps we can become better at recognising patterns worth carrying forward.

And then we can help them find each other.


9. There's a catch: we may have to loosen our grip on the outcome

There is something uncomfortable about all of this.

If I plant a seed because I need my seed to become the tree, complexity is likely to disappoint me.

Most experiments won't transform systems.

Some good ideas will disappear. Some will be absorbed into existing institutions and lose what made them interesting. Some will mutate. Someone else may take an idea and do something better with it.

And something that looks like failure today might unexpectedly become useful years from now.

That makes this difficult for those of us trained to specify outcomes and demonstrate contribution towards them.

Perhaps sometimes success needs to mean something more modest:

We introduced a pattern we thought was worth having in the system.

We learned from it.

We allowed other people to change it.

We connected it to other things.

And we paid attention to what happened next.


10. Finding the other weavers

I am still worried about AI.

And I'm worried about institutional decline, social fragmentation and some of the transitions happening in development and philanthropy.

Systems thinking doesn't give me reassurance that everything will work out.

But it does give me an alternative to the choice between controlling the future and being powerless before it.

We have agency without having control.

We can seed useful patterns.

We can protect experimentation.

We can strengthen institutions capable of learning.

We can notice when something interesting is emerging.

We can help useful patterns travel.

And, perhaps most importantly, we can connect.

Because if Patton is right about interconnectedness momentum, and transition theory is right about the importance of alternatives that already exist when systems begin to shift, then connection isn't something we do after we've figured out the solution.

Connection may itself be part of how the transition happens.

We don't know which seeds will grow.

We don't know what the system will select.

We don't know what South Africa — or the world — will look like on the other side of the transitions we are living through.

But we do have some agency over what we put into the system now.

Perhaps step one is finding the other weavers.

Image Source: https://indlulamithi.org.za/

A note on AI: There is a certain irony in writing this post with AI. I used ChatGPT as a thinking and drafting partner: to test the argument, connect it to theory, find sources, and help shape the writing. The ideas and examples I wanted to explore were mine; I made the substantive choices and checked the sources. I remain responsible for what is written here.

Thursday, September 03, 2026

South Africa's National Evaluation Policy Framework 2025

 Ooh — I missed this.

The National Evaluation Policy Framework 2025, published in February 2026, is worth a read for anyone doing evaluation work in South Africa. Not necessarily because it will revolutionise how we evaluate, but because it gives a useful sense of the definitions, questions and forms of evaluative practice that our national system, and the Department of Planning, Monitoring and Evaluation, considers important.

There are some things I particularly like.

First, the framework goes beyond simply listing the familiar evaluation types — diagnostic, design, implementation, outcome, economic, impact and synthesis evaluations. It also deliberately broadens the conversation to evaluative practices and approaches: evaluative thinking in monitoring, evaluative workshops, rapid evaluations, evidence synthesis, institutional reviews, systems thinking, foresight and equity considerations.

I like this. Evaluation is not only something that happens when we commission a formal evaluation. Evaluative thinking can be embedded in monitoring, reflection, strategy and decision-making. And I was particularly pleased to see systems thinking in evaluation explicitly named.

The sustainability criterion is also nicely expansive. It asks the deceptively simple question: “Will the benefits last?” But the definition goes much further than financial sustainability, asking us to consider the economic, social, environmental and institutional capacities of the systems needed to sustain benefits, as well as resilience, risks and trade-offs.

And then there are two criteria we should be rather proud of: Climate and Ecosystems Health (CEH) and Transformative Equity (TE). These are South African additions to the usual OECD-DAC criteria, explicitly bringing environmental systems and systemic inequity into how we judge interventions.

That feels like a genuinely useful contribution from South Africa to evaluation practice.

Source: Based on the National Evaluation Policy Framework 2025 (DPME, 2026). Concept and content selection by Benita Williams; infographic created with ChatGPT/OpenAI.

Thursday, August 27, 2026

The word "Sustainability" is a tricky one

I have been thinking about sustainability for a very long time. Long enough, in fact, to write a Master's thesis about it. 

And I think the reason I keep coming back to it is that sustainability sounds deceptively simple. We usually ask something like: Will the benefits last after the programme ends? That is essentially how the OECD DAC frames it too: do the net benefits continue, or are they likely to continue? The newer definition is actually quite useful because it also asks us to think about the financial, social, environmental and institutional capacities needed to keep those benefits going.

But hang on...What exactly is it that we expect to continue? The programme? The activities? The funding? The people? The materials? The way of working? The result? Those are not the same thing.

The sustainability literature makes this distinction too. Sometimes the intervention itself survives. Sometimes particular activities continue. Sometimes a function becomes institutionalised somewhere else. And sometimes the original programme disappears almost completely, but the benefit continues.

I am increasingly less interested in whether the programme survives, and more interested in whether the change survives. Or perhaps survives is the wrong word as well.

What if sustainability involves changing?

This is where systems thinking became useful to me.

In my Master's research I used Stockmann's idea of extended dynamic sustainability. The basic idea is that an intervention enters an existing system. It interacts with people, organisations, rules, relationships, other programmes and whatever else is happening at the time. Then the system keeps moving. So the outcomes we see ten years later do not necessarily look like the outcomes we saw at endline. They may have changed shape completely.

The opposite assumption is slightly odd when you think about it. We introduce an intervention into a changing education system, leave, come back five or ten years later, and then ask whether the thing is still exactly where we left it. Why would it be?

People move. Policies change. Organisations merge. New programmes arrive. Funding disappears. Somebody adapts a tool because the original version no longer works. Someone who was trained moves into a different job and takes the idea with them. Some of that might actually be evidence of sustainability.

In the education case I studied for my Master's, three mechanisms seemed particularly useful for making sense of this.

One was problem-solving. Did the intervention leave people or institutions better able to solve the problem themselves? Another was multiplication. Did the benefit spread beyond the original people or places? And then modelling: did somebody else take the idea, process or structure and adapt it as a useful model?

That work made me think that some of the strongest forms of sustainability may be almost invisible if we only look for the original intervention. A programme may be gone. But the system may have learned something.

Which brings me to the Anglo American Education programme

I was lucky enough to work on the evaluation of this programme, alongside colleagues, over a very long period — from 2019 to 2026. Long enough to see different phases of the programme, different implementation models, different schools and Early Childhood Development (ECD) centres, and eventually to go back and ask what had actually lasted.

The programme itself was ambitious. It worked across schools and ECD centres in communities around Anglo American operations, and it tried to strengthen several parts of the education system at once. In the school programme this included teacher development, school leadership, parent engagement, learner support, infrastructure and, later, ICT. In ECD, the support included infrastructure, learning materials, practitioner and centre-manager training and coaching, and support with registration and compliance.

So this was not a neat little intervention where one activity was expected to produce one outcome. And perhaps more importantly for this discussion, it was not a programme that forgot about sustainability and suddenly discovered it in the final six months. Quite the opposite.

The programme thought very carefully about stakeholder engagement, ownership and government alignment. Engagement started before implementation and continued throughout the programme. It involved national, provincial, district, community, school and ECD-level actors. There were Local Management Committees. There was provincial project-management capacity. Issues were monitored. The programme deliberately tried to respond to local context rather than simply roll out a fixed package.

In Phase 2, schools were even asked to opt in rather than simply being selected, precisely because the designers thought greater ownership might matter. So yes, this was a pretty thoughtful design. Which makes what happened later much more interesting.

Because then we went back.

Quite a lot was still there

In ECD, the initial results had been strong. The programme combined infrastructure and materials with practitioner training and coaching, support to ECD managers, and help with registration and compliance. And the child outcomes were striking. The proportion of assessed children who were developmentally on track increased from 38% to 65%, moving from below the national comparison benchmark to above it. When we returned later, many centres were still operating. They were still serving children. Many had retained registration, were receiving subsidies and were still using structured learning materials.

In the Whole School Development work, some of the strongest evidence of sustainment was surprisingly ordinary: reading routines, phonics practices, planning approaches and ways of supporting struggling learners were still being used. And I think there is something in that. Maybe things become sustainable when they stop feeling like programme activities and simply become useful ways of doing the work.

But context kept interfering

Some ECD centres were struggling with delayed subsidy payments. Some were navigating registration changes. Some had become overcrowded. Security was a problem. And this is where I think we sometimes use the word context too casually. We write something like: “The intervention was implemented in a challenging context.” Then we move on. But the context is doing quite a lot of work there.

A registered ECD centre with stable management, enough practitioners and reliable subsidy payments is not receiving the same intervention as a centre with an acting manager, overcrowded classrooms and late payments. The input may be identical. The system around it isn't. And that matters for what gets sustained.

So what does this mean for ECD investment now?

I think this is where the systems view becomes practical rather than philosophical. South African ECD now has the benefit of a more clearly negotiated system direction through initiatives such as the Bana Pele Blueprint. That does not mean every funder should suddenly run the same programme. Please no.

Different organisations should still do different things. One may focus on infrastructure. Another on practitioner quality. Another on registration. Another on nutrition or financing or data. But it helps enormously if they are contributing towards a shared system goal. So perhaps an early programme-design question should be:

Where does this intervention fit in the wider ECD system?

And then:

What problem in that system are we helping to solve? Who else is already working on it? Which existing structures could carry this work later? Are we strengthening them, or quietly building a parallel system next door? And if our particular project disappears in five years' time, what capability will remain? 

I no longer think sustainability means freezing an intervention in place. 

I think it may be much more about what the system has learned to do because the intervention was there. Maybe the programme disappears. Maybe the materials change. Maybe somebody else takes over the function. Maybe the original model gets adapted beyond recognition. If the system is now better able to solve the problem, is this not sustainability?  

Photo by Alina Grubnyak on Unsplash

Wednesday, August 26, 2026

Someone made me study philosophy of science and now I’m triggered


I cringe a little every time someone talks about creating “one source of truth”.

It happened yesterday at the NASCEE seminar, in a discussion about NED Connect. I heard something similar at Jon Molver’s PULSE event.  More troublingly, a government official used almost exactly the same language when talking about COGTA’s National Strategic Hub. I thought I might be overreacting — until the COGTA slide came up. It literally said: “Our single source of truth.”

Each time, my immediate reaction is: Orwellian. Ministry of Truth, anyone? I know, I know — that is probably unfair.

What people probably mean is something much more practical: Can we please agree on which dataset we are using? They are talking about getting away from multiple spreadsheets, competing numbers, inconsistent definitions and records that nobody quite knows whether to trust. I am completely sympathetic to that problem.

People use words without necessarily thinking through everything they imply. I certainly do. But in the intellectual tradition in which I work, words like truth, evidence, validity and objectivity are triggers. We have spent rather a lot of time arguing about what they mean. So I cannot quite hear “one source of truth” innocently.

A small philosophy-of-science detour

There is a much longer investigation here about my own worldview that I am not going to inflict on you now. The footnote version is this: I also twitch when somebody says “science has proved that…”. In the Popperian tradition I was taught, science is not really in the business of proving things true. It advances by making claims that can, in principle, be shown to be wrong. Theories that survive repeated attempts to falsify them become more credible, but they do not become sacred. That little distinction probably explains quite a lot about why “one source of truth” bothers me.

Science becomes trustworthy not because nobody can challenge it, but because people can. It becomes trustworthy because people can challenge it. 

When a data system is described as the source of truth, I hear something rather more final than I suspect the speaker intends.

And then there is power

This is where my evaluator brain kicks in. Robert Picciotto puts it rather starkly:

“Evaluation is not value-free. It is steeped in politics.”

He also asks whose goals matter, which values are being used to judge merit and worth, and who gains or loses from the methods we choose. Those questions travel quite easily into the world of data systems.

Who decided the categories? What becomes easy to see because it fits them? What becomes harder to see because it does not? Which questions was the system designed to answer in the first place? Every database has to make choices like this. The difficulty comes when the choices disappear from view.

If we map an ecosystem, for example, we may decide to classify organisations by geography, programme area, funding, reach or organisational type. Those may be entirely sensible choices. But over time the structure we designed can start to feel like the natural structure of the sector itself. What is easy to count becomes easier to discuss, and what does not fit neatly into the categories can fade into the background. Perhaps this is where the rather grand word hegemony becomes useful. One way of seeing the world can become so familiar that we stop noticing it is only one way of seeing. 

The danger is that one representation of reality acquires the authority of reality itself. That is the bit that makes me nervous about the word truth

This is hardly a new anxiety. In African evaluation, the Made in Africa Evaluation movement has been asking related questions for years about whose knowledge counts, whose categories shape the field, and what happens when one knowledge tradition comes to look universal. Zenda Ofir and Adeline Sibanda’s chapter, Made in Africa Evaluation: Decolonizing the Past, Present, and Future, is a useful entry point.

Somebody still has to make sense of it

There is another reason I am reluctant to hand truth over to the database: Data do not interpret themselves.

Even a very good system still needs someone who understands where the information came from, what the measures can actually tell us, and where apparently comparable numbers are not quite comparable after all. Context matters too. A pattern in a dashboard may be significant. It may also be an artefact of how something was classified, when the data were collected, or who reported it.

The interesting work starts once the data have been assembled. Somebody has to move between the numbers and the real world they are supposed to describe. Somebody has to turn information into a useful account without making it sound more certain than it is.

This is one reason I am not especially worried that dashboards, integrated platforms or AI are going to put evaluators out of work. They may change the work considerably. I hope they do.

If evaluators remain useful, I suspect it will be because we are comfortable occupying that slightly uncomfortable space between evidence and decision-making. We know enough about methods to be cautious about what a number can actually tell us, and enough about context to know when the number is probably not telling the whole story.

And, ideally, we can explain what we see in a way that somebody else can understand.

So who should do the sensemaking?

This is where I make the case for having people around who understand more than the mechanics of data.

You probably want someone who understands the data itself, certainly. But I also want someone who has thought a little about how knowledge gets made, how we decide whether a claim is credible, and when we may be claiming more than the evidence supports.

That may be an evaluator. It may be a researcher, an analyst or somebody else entirely. The label matters less to me than the way of thinking.

Can this person work carefully with evidence without ignoring uncertainty? Can they notice when a clean representation has left something important out or when it is oversimplifying? That is the kind of sensemaking I want sitting next to a powerful data system.

Truth, Beauty and Justice

This brings me back to Ernest House and one of my favourite ideas in evaluation: Truth, Beauty and Justice. I like the formulation because it gives me a much better way to think about what good sensemaking should aspire to.

Truth asks whether the claims we are making can be defended.

Beauty is about whether we can make sense of the evidence and communicate it in a way that helps people see something more clearly.

Justice makes us pay attention to whose experience is represented in that account, and whose may have been left out.

That is probably where I land on all of this. I want the good database. I want the messy spreadsheets sorted out. I want shared definitions and better systems. I just do not want the system itself to become the truth.

Give me good data, and then give me someone who understands what it can and cannot tell us. Make sure they appreciate what it means to make a claim about truth.  Can they also please have the ability to translate the evidence beautifully and care about justice

Evaluators may not be out of a job just yet. (At least, not if that is the job we think we are here to do).

Photo by Abdul Ahad Sheikh on Unsplash