Thursday, September 03, 2026

South Africa's National Evaluation Policy Framework 2025

 Ooh — I missed this.

The National Evaluation Policy Framework 2025, published in February 2026, is worth a read for anyone doing evaluation work in South Africa. Not necessarily because it will revolutionise how we evaluate, but because it gives a useful sense of the definitions, questions and forms of evaluative practice that our national system, and the Department of Planning, Monitoring and Evaluation, considers important.

There are some things I particularly like.

First, the framework goes beyond simply listing the familiar evaluation types — diagnostic, design, implementation, outcome, economic, impact and synthesis evaluations. It also deliberately broadens the conversation to evaluative practices and approaches: evaluative thinking in monitoring, evaluative workshops, rapid evaluations, evidence synthesis, institutional reviews, systems thinking, foresight and equity considerations.

I like this. Evaluation is not only something that happens when we commission a formal evaluation. Evaluative thinking can be embedded in monitoring, reflection, strategy and decision-making. And I was particularly pleased to see systems thinking in evaluation explicitly named.

The sustainability criterion is also nicely expansive. It asks the deceptively simple question: “Will the benefits last?” But the definition goes much further than financial sustainability, asking us to consider the economic, social, environmental and institutional capacities of the systems needed to sustain benefits, as well as resilience, risks and trade-offs.

And then there are two criteria we should be rather proud of: Climate and Ecosystems Health (CEH) and Transformative Equity (TE). These are South African additions to the usual OECD-DAC criteria, explicitly bringing environmental systems and systemic inequity into how we judge interventions.

That feels like a genuinely useful contribution from South Africa to evaluation practice.

Source: Based on the National Evaluation Policy Framework 2025 (DPME, 2026). Concept and content selection by Benita Williams; infographic created with ChatGPT/OpenAI.

Thursday, August 27, 2026

The word "Sustainability" is a tricky one

I have been thinking about sustainability for a very long time. Long enough, in fact, to write a Master's thesis about it. 

And I think the reason I keep coming back to it is that sustainability sounds deceptively simple. We usually ask something like: Will the benefits last after the programme ends? That is essentially how the OECD DAC frames it too: do the net benefits continue, or are they likely to continue? The newer definition is actually quite useful because it also asks us to think about the financial, social, environmental and institutional capacities needed to keep those benefits going.

But hang on...What exactly is it that we expect to continue? The programme? The activities? The funding? The people? The materials? The way of working? The result? Those are not the same thing.

The sustainability literature makes this distinction too. Sometimes the intervention itself survives. Sometimes particular activities continue. Sometimes a function becomes institutionalised somewhere else. And sometimes the original programme disappears almost completely, but the benefit continues.

I am increasingly less interested in whether the programme survives, and more interested in whether the change survives. Or perhaps survives is the wrong word as well.

What if sustainability involves changing?

This is where systems thinking became useful to me.

In my Master's research I used Stockmann's idea of extended dynamic sustainability. The basic idea is that an intervention enters an existing system. It interacts with people, organisations, rules, relationships, other programmes and whatever else is happening at the time. Then the system keeps moving. So the outcomes we see ten years later do not necessarily look like the outcomes we saw at endline. They may have changed shape completely.

The opposite assumption is slightly odd when you think about it. We introduce an intervention into a changing education system, leave, come back five or ten years later, and then ask whether the thing is still exactly where we left it. Why would it be?

People move. Policies change. Organisations merge. New programmes arrive. Funding disappears. Somebody adapts a tool because the original version no longer works. Someone who was trained moves into a different job and takes the idea with them. Some of that might actually be evidence of sustainability.

In the education case I studied for my Master's, three mechanisms seemed particularly useful for making sense of this.

One was problem-solving. Did the intervention leave people or institutions better able to solve the problem themselves? Another was multiplication. Did the benefit spread beyond the original people or places? And then modelling: did somebody else take the idea, process or structure and adapt it as a useful model?

That work made me think that some of the strongest forms of sustainability may be almost invisible if we only look for the original intervention. A programme may be gone. But the system may have learned something.

Which brings me to the Anglo American Education programme

I was lucky enough to work on the evaluation of this programme, alongside colleagues, over a very long period — from 2019 to 2026. Long enough to see different phases of the programme, different implementation models, different schools and Early Childhood Development (ECD) centres, and eventually to go back and ask what had actually lasted.

The programme itself was ambitious. It worked across schools and ECD centres in communities around Anglo American operations, and it tried to strengthen several parts of the education system at once. In the school programme this included teacher development, school leadership, parent engagement, learner support, infrastructure and, later, ICT. In ECD, the support included infrastructure, learning materials, practitioner and centre-manager training and coaching, and support with registration and compliance.

So this was not a neat little intervention where one activity was expected to produce one outcome. And perhaps more importantly for this discussion, it was not a programme that forgot about sustainability and suddenly discovered it in the final six months. Quite the opposite.

The programme thought very carefully about stakeholder engagement, ownership and government alignment. Engagement started before implementation and continued throughout the programme. It involved national, provincial, district, community, school and ECD-level actors. There were Local Management Committees. There was provincial project-management capacity. Issues were monitored. The programme deliberately tried to respond to local context rather than simply roll out a fixed package.

In Phase 2, schools were even asked to opt in rather than simply being selected, precisely because the designers thought greater ownership might matter. So yes, this was a pretty thoughtful design. Which makes what happened later much more interesting.

Because then we went back.

Quite a lot was still there

In ECD, the initial results had been strong. The programme combined infrastructure and materials with practitioner training and coaching, support to ECD managers, and help with registration and compliance. And the child outcomes were striking. The proportion of assessed children who were developmentally on track increased from 38% to 65%, moving from below the national comparison benchmark to above it. When we returned later, many centres were still operating. They were still serving children. Many had retained registration, were receiving subsidies and were still using structured learning materials.

In the Whole School Development work, some of the strongest evidence of sustainment was surprisingly ordinary: reading routines, phonics practices, planning approaches and ways of supporting struggling learners were still being used. And I think there is something in that. Maybe things become sustainable when they stop feeling like programme activities and simply become useful ways of doing the work.

But context kept interfering

Some ECD centres were struggling with delayed subsidy payments. Some were navigating registration changes. Some had become overcrowded. Security was a problem. And this is where I think we sometimes use the word context too casually. We write something like: “The intervention was implemented in a challenging context.” Then we move on. But the context is doing quite a lot of work there.

A registered ECD centre with stable management, enough practitioners and reliable subsidy payments is not receiving the same intervention as a centre with an acting manager, overcrowded classrooms and late payments. The input may be identical. The system around it isn't. And that matters for what gets sustained.

So what does this mean for ECD investment now?

I think this is where the systems view becomes practical rather than philosophical. South African ECD now has the benefit of a more clearly negotiated system direction through initiatives such as the Bana Pele Blueprint. That does not mean every funder should suddenly run the same programme. Please no.

Different organisations should still do different things. One may focus on infrastructure. Another on practitioner quality. Another on registration. Another on nutrition or financing or data. But it helps enormously if they are contributing towards a shared system goal. So perhaps an early programme-design question should be:

Where does this intervention fit in the wider ECD system?

And then:

What problem in that system are we helping to solve? Who else is already working on it? Which existing structures could carry this work later? Are we strengthening them, or quietly building a parallel system next door? And if our particular project disappears in five years' time, what capability will remain? 

I no longer think sustainability means freezing an intervention in place. 

I think it may be much more about what the system has learned to do because the intervention was there. Maybe the programme disappears. Maybe the materials change. Maybe somebody else takes over the function. Maybe the original model gets adapted beyond recognition. If the system is now better able to solve the problem, is this not sustainability?  

Photo by Alina Grubnyak on Unsplash

Wednesday, August 26, 2026

Someone made me study philosophy of science and now I’m triggered


I cringe a little every time someone talks about creating “one source of truth”.

It happened yesterday at the NASCEE seminar, in a discussion about NED Connect. I heard something similar at Jon Molver’s PULSE event.  More troublingly, a government official used almost exactly the same language when talking about COGTA’s National Strategic Hub. I thought I might be overreacting — until the COGTA slide came up. It literally said: “Our single source of truth.”

Each time, my immediate reaction is: Orwellian. Ministry of Truth, anyone? I know, I know — that is probably unfair.

What people probably mean is something much more practical: Can we please agree on which dataset we are using? They are talking about getting away from multiple spreadsheets, competing numbers, inconsistent definitions and records that nobody quite knows whether to trust. I am completely sympathetic to that problem.

People use words without necessarily thinking through everything they imply. I certainly do. But in the intellectual tradition in which I work, words like truth, evidence, validity and objectivity are triggers. We have spent rather a lot of time arguing about what they mean. So I cannot quite hear “one source of truth” innocently.

A small philosophy-of-science detour

There is a much longer investigation here about my own worldview that I am not going to inflict on you now. The footnote version is this: I also twitch when somebody says “science has proved that…”. In the Popperian tradition I was taught, science is not really in the business of proving things true. It advances by making claims that can, in principle, be shown to be wrong. Theories that survive repeated attempts to falsify them become more credible, but they do not become sacred. That little distinction probably explains quite a lot about why “one source of truth” bothers me.

Science becomes trustworthy not because nobody can challenge it, but because people can. It becomes trustworthy because people can challenge it. 

When a data system is described as the source of truth, I hear something rather more final than I suspect the speaker intends.

And then there is power

This is where my evaluator brain kicks in. Robert Picciotto puts it rather starkly:

“Evaluation is not value-free. It is steeped in politics.”

He also asks whose goals matter, which values are being used to judge merit and worth, and who gains or loses from the methods we choose. Those questions travel quite easily into the world of data systems.

Who decided the categories? What becomes easy to see because it fits them? What becomes harder to see because it does not? Which questions was the system designed to answer in the first place? Every database has to make choices like this. The difficulty comes when the choices disappear from view.

If we map an ecosystem, for example, we may decide to classify organisations by geography, programme area, funding, reach or organisational type. Those may be entirely sensible choices. But over time the structure we designed can start to feel like the natural structure of the sector itself. What is easy to count becomes easier to discuss, and what does not fit neatly into the categories can fade into the background. Perhaps this is where the rather grand word hegemony becomes useful. One way of seeing the world can become so familiar that we stop noticing it is only one way of seeing. 

The danger is that one representation of reality acquires the authority of reality itself. That is the bit that makes me nervous about the word truth

This is hardly a new anxiety. In African evaluation, the Made in Africa Evaluation movement has been asking related questions for years about whose knowledge counts, whose categories shape the field, and what happens when one knowledge tradition comes to look universal. Zenda Ofir and Adeline Sibanda’s chapter, Made in Africa Evaluation: Decolonizing the Past, Present, and Future, is a useful entry point.

Somebody still has to make sense of it

There is another reason I am reluctant to hand truth over to the database: Data do not interpret themselves.

Even a very good system still needs someone who understands where the information came from, what the measures can actually tell us, and where apparently comparable numbers are not quite comparable after all. Context matters too. A pattern in a dashboard may be significant. It may also be an artefact of how something was classified, when the data were collected, or who reported it.

The interesting work starts once the data have been assembled. Somebody has to move between the numbers and the real world they are supposed to describe. Somebody has to turn information into a useful account without making it sound more certain than it is.

This is one reason I am not especially worried that dashboards, integrated platforms or AI are going to put evaluators out of work. They may change the work considerably. I hope they do.

If evaluators remain useful, I suspect it will be because we are comfortable occupying that slightly uncomfortable space between evidence and decision-making. We know enough about methods to be cautious about what a number can actually tell us, and enough about context to know when the number is probably not telling the whole story.

And, ideally, we can explain what we see in a way that somebody else can understand.

So who should do the sensemaking?

This is where I make the case for having people around who understand more than the mechanics of data.

You probably want someone who understands the data itself, certainly. But I also want someone who has thought a little about how knowledge gets made, how we decide whether a claim is credible, and when we may be claiming more than the evidence supports.

That may be an evaluator. It may be a researcher, an analyst or somebody else entirely. The label matters less to me than the way of thinking.

Can this person work carefully with evidence without ignoring uncertainty? Can they notice when a clean representation has left something important out or when it is oversimplifying? That is the kind of sensemaking I want sitting next to a powerful data system.

Truth, Beauty and Justice

This brings me back to Ernest House and one of my favourite ideas in evaluation: Truth, Beauty and Justice. I like the formulation because it gives me a much better way to think about what good sensemaking should aspire to.

Truth asks whether the claims we are making can be defended.

Beauty is about whether we can make sense of the evidence and communicate it in a way that helps people see something more clearly.

Justice makes us pay attention to whose experience is represented in that account, and whose may have been left out.

That is probably where I land on all of this. I want the good database. I want the messy spreadsheets sorted out. I want shared definitions and better systems. I just do not want the system itself to become the truth.

Give me good data, and then give me someone who understands what it can and cannot tell us. Make sure they appreciate what it means to make a claim about truth.  Can they also please have the ability to translate the evidence beautifully and care about justice

Evaluators may not be out of a job just yet. (At least, not if that is the job we think we are here to do).

Photo by Abdul Ahad Sheikh on Unsplash

Parenting in the age of AI

Our friends at NASCEE recently sent out an invitation to the first seminar on AI in Basic Education, as part of the Lefa AI initiative.

As an evaluator, I am extremely interested in this topic. As a parent of school-going children, I am perhaps even more fascinated by it. Although “fascinated” may not be quite the right word. It is the kind of fascination where you lean in because you really want to see what happens next, while at the same time keeping one hand firmly over your heart in case what happens next gives it a shock.

AI in education raises all sorts of practical questions about teaching, assessment, access and learning. But in our house, some of the questions already feel much more immediate. What should children still learn to do for themselves? When does using AI become a useful extension of their abilities, and when does it quietly remove the struggle through which those abilities are developed? And perhaps most difficult of all: how do we explain to children why they should spend time learning something when a machine can produce a polished version in seconds?

I do not have tidy answers. But watching my own children use these tools has made the questions much harder to ignore.

A few days ago, my youngest showed me an English dialogue she had written for school, and I was genuinely surprised. It was funny. The characters had personalities. The dialogue moved. There was timing and a sense of scene. And she had done it on her own.

I would love to say this came from years of reading. Heavens knows it did not. Reading is an ongoing struggle in our house. My best guess is that she has absorbed enough stories through television and other media to understand, almost intuitively, how characters behave and how dialogue works. That in itself is interesting. We often speak about creativity as though it emerges from nowhere, but of course it does not. Human beings are pattern learners. We absorb stories, rhythms, phrases, gestures and conventions from the culture around us, and then we recombine them into something that feels like our own.

My son presents almost the opposite problem. Writing does not come easily to him, and he is very willing to offload cognitive work. But he is also quite inventive about it. For his school entrepreneur day, he decided he needed a Wheel of Fortune-style spinner for a game, so he used ChatGPT to programme one. He also gets somebody else to prepare accurate study notes, then feeds those notes into an AI system that turns them into flashcards, memory prompts and other study tools.

Part of me thinks: clever. Another part of me thinks: oh dear. Because he may become very good at orchestrating tools without ever becoming very good at writing. What does the future look like? Must he necessarily be good at writing? 

And I am beginning to think that this is where the AI debate becomes much harder when we are talking about children. As adults, cognitive offloading is often sensible. I use a calculator because I already understand arithmetic. I use spellcheck because I can spell well enough to know when it is wrong. I can ask AI to tighten a paragraph because I have spent decades learning how paragraphs work.

But what happens when we offload the work before the underlying capability has been built?

A child learns to write by writing bad sentences, then slightly better ones, then awkward paragraphs, and eventually paragraphs that begin to say what they intended. The inefficiency is not a bug. It is the learning process. AI can remove a great deal of that struggle, which sounds wonderful until we remember that the struggle may be exactly what develops the capability.

So the question is not simply whether my son is “cheating”. It is more uncomfortable than that. What experiences does he still need to have in order to become someone who can think and write for himself? If AI consistently occupies that space before the skill is established, he may become very good at directing writing without ever becoming very good at writing. Those are not the same thing.

And then there is my daughter. What happens when she spends two hours creating something funny and original, only to discover that somebody else can type a prompt into an LLM and produce something equally polished in three seconds? I worry not only about skill loss, but about motivation. Why bother mastering something when the output itself can be reproduced almost instantly?

That question troubles me because creativity is not valuable only because it produces a poem, photograph, song or school dialogue. There is value in becoming the person who can make it. The capability matters. The eye matters. The judgement matters. The satisfaction of being able to do something matters.

She is also becoming interested in photography, and perhaps one day AI will extend what she can create. I do not necessarily object to that. If she develops an eye for composition, light, framing and emotion, AI may become an expressive tool. But if the technology does the aesthetic work before she has had the chance to develop those instincts herself, then something else may be happening.

I increasingly find myself returning to one question:

What human capability is the technology extending, and what capability is it bypassing?

Perhaps that is the distinction I need as a parent. AI after mastery may be an extraordinary tool. AI before mastery may sometimes quietly remove the very practice through which mastery develops.

And unfortunately, “because you still have to learn it first” is going to be a much harder argument to make to our children than it ever was to us.

What if you spent five days getting the learning architecture right at the start?

 


What if you spent five days getting the learning architecture right at the start?

Have you ever developed a project proposal with a slight sense of dread?

You can have a perfectly sensible intervention. A clear set of deliverables. A committed implementation team. You can even deliver exactly what you promised. And still find that very little shifts because the surrounding system does not move with it. I keep coming back to that.

We have become much more interested in systems thinking in the development sector, but I wonder whether we still bring it in too late. We talk about systems when we evaluate programmes, or when we are trying to understand why outcomes did or did not happen. Maybe we should be doing more of that thinking right at the beginning.

The same goes for Theory of Change work. We have known for years that participatory approaches matter in evaluation. When implementers sit together and work through the logic of a programme, the value is not only the diagram they produce. Quite often, the useful bit is the conversation that happens before everyone agrees on the arrows.

Bring the funder into that conversation and it gets even more interesting. You start surfacing expectations, assumptions and constraints on both sides. Sometimes people discover that they have been using the same words to mean slightly different things.

The process matters as much as the diagram.

A Theory of Change is much more useful when the diagram is the record of a good conversation, rather than the purpose of the exercise. And once that conversation gets going, the questions tend to change.

  • What problem are we actually trying to change?
  • Who else is working on it, and what have they learned?
  • What do we think will make a difference, and what has to be true for that to happen?
  • What should we pay attention to while implementation unfolds?
  • If the programme hopes to influence a wider system, what would we expect to see changing beyond the immediate project results?

And perhaps the question I like most:

What are we genuinely trying to learn?

That is where the Theory of Change starts becoming something more useful. It becomes the beginning of a learning architecture:

Theory of Change → indicators and evidence → learning questions → reflection and reporting → evaluation → adaptation and future decisions.

So how might we do this differently?

One approach we are increasingly interested in is to spend a small amount of concentrated time up front with people who understand the sector, can work practically with indicators and evaluation design, and are willing to ask the difficult questions.

Not only:

“How will we measure whether this worked?”

But sometimes also:

“Should we really be doing this?”

We think five focused days can create useful space for that work. Not five consecutive days sitting in a workshop room. Some of it is better face to face. Some can happen in focused online sessions. And some is technical work between engagements: reading documents, testing the logic, refining indicators, and working the emerging thinking back into the Theory of Change and Learning Agenda.

Days 1–2: Understand what we are dealing with

The first part is really about getting to know the territory. And by territory, I do not only mean the intervention. We want to understand the organisation around it too. What is the strategy? How is long-term progress already being tracked? What does management need to know? What does the Board expect to see? What are the contractual reporting requirements? How strong is the internal M&E capacity? What is the reporting rhythm? How does the organisation actually learn?

Those things matter more than they sometimes appear to. You can design a beautiful framework, but if nobody has the time or systems to collect the information, it is not going to help much. We also want to understand what the organisation means when it talks about systemic change.

  • Is the programme mainly expected to deliver a defined set of outputs and outcomes?
  • Or is there also an expectation that it should influence relationships, institutions, practice, policy, funding decisions or the behaviour of other actors?

Then we look outside the programme.

  • Who else is working on the same problem?
  • What are government, funders or civil society organisations already doing?
  • Where might there be duplication?
  • Where are the gaps?
  • Is there somebody the programme should probably be talking to?

We are not pretending that two days gives you a systems map. It doesn’t. But you can learn quite a lot by simply asking who else is working on the problem and what they already know. Sometimes that changes the programme logic quite substantially.

Days 3–5: Work on the Theory of Change, measures and Learning Agenda

Then we get into the technical work. We use whatever already exists. That might be a Theory of Change diagram, a logframe, an indicator matrix, a proposal, a strategy document, a contract — or several of these that do not quite agree with each other. That happens more often than one might hope. The point is not to replace everything. It is to work out whether the different pieces tell a coherent story.

  • What is the intended change?
  • How is the programme supposed to contribute to it?
  • What assumptions are we making?
  • What might get in the way?
  • Which parts of the logic are reasonably within the programme’s control, and which depend on other actors or on the wider system?

This sounds like quite a technical distinction, but it changes the conversation a lot.

  • What are we actually accountable for?
  • What are we hoping to influence?
  • And what are we still genuinely uncertain about?

I think those are much better questions to sort out at the beginning than halfway through an evaluation. Then comes the other very practical question:

What can we actually measure and report on?

Not what would be lovely to know if we had unlimited time and data.

  • What can this programme realistically track?
  • Which indicators are meaningful?
  • Which are feasible?
  • Which ones will somebody actually use?
  • And which ones are we collecting because they have somehow always been there?

From there, we can start building the Learning Agenda.

  • Some information is needed to manage implementation.
  • Some is required for formal reporting.
  • Some matters because management or the Board needs it.
  • And some questions are worth returning to simply because we do not know the answer yet.

That last category is important. Not every question needs an indicator. Not every indicator needs to be discussed every month. Not every interesting issue needs an evaluation. And not everything worth learning can be reduced to a number. We also need to decide how the learning will actually happen.

  • What gets discussed regularly?
  • What belongs in a formal report?
  • What should trigger a deeper conversation?
  • Where would a deliberate Learning Moment help?
  • What might need a more rigorous evaluation later?
  • And how will what we learn actually feed back into decisions?

That, for me, is the point.

The result should not be another MEL document that sits on a shared drive. It should be a shared understanding of what the programme is trying to change, what it can reasonably be held accountable for, what sits outside its control, and what everybody wants to keep learning while the work unfolds.

Five days will obviously not solve everything. But I suspect five good days at the beginning can save a lot of confusion later. And perhaps more importantly, they can get some of the difficult questions onto the table before the programme architecture hardens around them.

Friday, July 03, 2026

What Guba and Lincoln offer to the use of AI debate

“Thou shalt use AI responsibly”





The current debate about AI often gets stuck in a fairly familiar place.

You can use AI, but you need to use it “well”.
Do not pass AI-generated work off as your own.
Do not let it do your analysis for you.
Do not let it write your report for you.
Be careful: it hallucinates references.

All of this is true. But I think it misses something more subtle.

AI, chocolate and cognitive offloading

For many evaluators, researchers and knowledge workers, the real risk is not that we deliberately ask AI to do work we know we should be doing ourselves. The real risk is that, when we are tired, under pressure, cognitively overloaded or facing a difficult piece of judgement, we slowly let the tool carry more of the thinking than we intended.

The issue is cognitive offloading.

Cognitive offloading is not new. Humans have always used external tools to reduce mental effort: notebooks, calculators, checklists, templates, search engines, colleagues, diagrams and frameworks. In that sense, AI is part of a long history of tools that extend our thinking.

But generative AI changes the nature of the offloading. It does not only store information or speed up a calculation. It can summarise, classify, infer, draft, compare, frame and argue. In other words, it can begin to look like analysis.

And here is the uncomfortable part: even when we know the risks, we are still human.

For anyone who has ever tried to diet while sitting next to chocolate, this is not hard to understand. You may have a clear rule. You may genuinely believe in the rule. You may even have defended the rule publicly. But when you are tired, stressed or depleted, the chocolate becomes much harder to resist.

AI is a bit like that.

The concern is not only: Did you use AI?

The more important question is: At what point did the judgement move?

This is where I think Guba and Lincoln have something surprisingly useful to offer.

 

Guba and Lincoln enters the room….

In a meeting recently, someone made a throwaway comment about Guba and Lincoln. I responded with a small verbal protest. This was probably lost on anyone who has wisely avoided the old evaluation “paradigm wars” — including the Lincoln–Sechrest debate, where Lincoln stood for constructivist, stakeholder-shaped evaluation, and Sechrest defended disciplined evidence, validity and causal reasoning.

Niche evaluation lore? Definitely. 

Still useful for thinking about AI? Surprisingly, yes.

But afterwards I kept thinking about it. Guba and Lincoln’s Fourth Generation Evaluation challenged the idea that evaluation quality should be judged only by the conventional research criteria of validity, reliability, generalisability and objectivity. For constructivist, qualitative and naturalistic inquiry, they proposed a different language of quality: credibility, transferability, dependability and confirmability.

These criteria have since become familiar shorthand for trustworthiness in qualitative research. And trustworthiness, it seems to me, is exactly where the AI debate needs to go.


 From “Thou shalt not” to practical questions

Telling people not to offload their thinking is probably not enough.

It is a bit like saying:

Do not be biased.
Do not make weak inferences.
Do not overclaim.
Do not eat the chocolate.

The instruction is correct, but not very operational.

Guba and Lincoln help translate the concern into practical checks. Let me start with one of their criteria: dependability.

Dependability: is the process traceable?

The dependability criterion asks whether the inquiry process is coherent, logical and traceable.

In qualitative research, this usually means that the process should be documented well enough for someone else to understand how the findings were produced. It does not require mechanical replication, but it does require an audit trail: records of key decisions, changes in focus, coding choices, analytic memos, and the movement from raw data to interpretation.

For AI use, I find dependability especially useful because it helps me ask:

Is the human analytic process still visible, or has the route from evidence to claim become blurred?

Application 1: using AI to strengthen a coding framework

When I use AI in qualitative analysis, I do not start by asking it to “find the themes”. I begin by reading through the data myself, making notes on possible themes, codes, tensions and patterns.

Only then do I use AI as a thinking partner. For example, I might ask:

“Let’s develop a deductive coding framework. Based on the themes I have identified, suggest a coding memo for each theme, possible sub-themes, and five quotes I can use to test whether the code is working.”

At that point, the work becomes iterative. I might respond:

“I like the distinction you are making here, but I think these two codes overlap.”
“This code is too broad.”
“This seems to miss what participants are saying about implementation conditions.”
“Let’s test this theme against more quotes.”

I may then move into a more inductive round:

“I want to see what the data says about this theme. Pull together all quotes that seem related to it, including quotes that complicate or contradict it.”

This process does not make the coding framework dependable because AI produced it. It becomes more dependable because the framework has been tested, revised and documented through several rounds of analytic questioning.

AI helps speed up the iteration, but the analytic judgement remains mine.

Application 2: checking whether an AI-supported executive summary is dependable

I also use the dependability criterion when AI helps me draft or refine an executive summary.

An executive summary can become polished very quickly. That is useful, but also risky. The more fluent the summary becomes, the easier it is to lose sight of where each claim came from.

So I ask AI to help me rebuild the audit trail:

“Please create a table listing each claim in the executive summary. For each claim, indicate which document, section, dataset or finding it is based on.”

I then check the table against the source material. I correct weak links, remove claims that are not well supported, and adjust wording where the claim has become too strong.

I may then ask:

“Based on the information I provided, are there any findings, stakeholder views or pieces of evidence that might contradict or complicate these claims?”

This helps me check whether the summary has become too neat. If there are counterclaims or important qualifications, I decide whether the executive summary needs to acknowledge them, soften the wording, or adjust the argument.

Only after that do I use AI for board-ready refinement:

“Please simplify this into a clearer board-ready message.”
“Make the tone more direct.”
“Reduce repetition.”
“Sharpen the main implication.”

Again, the key issue is not whether AI was used. The key issue is whether I can still trace the movement from evidence, to claim, to final wording.

For me, that is what dependability looks like in an AI workflow: not avoiding AI, but using it in a way that keeps the analytic path visible.

And then, where it matters, I put the evidence trail into annexes.

Often many, many pages of annexes.

Will everyone read them? Almost certainly not. But that is not the only point. The annexes make the reasoning visible and verifiable. They show how claims were built, what evidence they rest on, and where the qualifications sit.

They also create a written record that future machine synthesis will probably love: structured, source-linked, explicit reasoning that can be checked, reused and interrogated.

So yes, the annexes may be unloved by human readers.

But they are part of the dependability infrastructure.

Beyond dependability

I have focused here on dependability because it is the criterion that most directly helps me manage AI use in my workflow. It asks whether the process is traceable — whether the path from evidence, to interpretation, to final wording is still visible.

But the same logic applies to the other trustworthiness criteria.

For credibility, I use AI to test whether claims are well supported, whether counter-evidence has been considered, and whether the final account is believable in relation to the data.

For transferability, I check whether the context has survived the synthesis: whether the specific conditions, boundaries and setting are still visible, rather than being smoothed into a generic lesson.

For confirmability, I ask whether I can still distinguish between what came from the evidence, what came from my judgement, and what AI helped me formulate.

The bottom line

So for me, Guba and Lincoln do not provide a nostalgic evaluation reference. They offer a practical way of asking a very current question:

Not “Did you use AI?” but “Was the use of AI trustworthy?”**

 And if the answer is yes, can you show how?

**************************************************************

**Yes, Charlie and I wrote this blogpost together - the giveaway ?

The "Not this, but that" structure (also known as contrastive negation or antithetical parallelism) is a highly recognizable rhetorical pattern heavily favored by ChatGPT and other LLMs.

*** Photo by June Gathercole on Unsplash



Thursday, March 21, 2024

A visualization on statistics choices

Another lovely resource - a decision tree to help choose the appropriate statistic to apply. 
And if you want to have step-by-step instructions on how to create it in SPSS, see this SAGE resource 





Wednesday, March 20, 2024

A visualisation of data visualization choices!

I always love it if people can simplify the complicated decisionmaking processes that we have in our heads, to a simple decision tree. A visualisation of data visualizationn choices!


 

 



Tuesday, March 05, 2024

Applying systems theoretical concepts to understand sustainability of education intervention outcomes

 

This Master’s dissertation addresses the research question: To what extent can the systems concept ‘extended dynamic sustainability’ be used to explain why some results of a donor-funded education development intervention were sustained ten years after its conclusion?



To address that question, the researcher identified a specific case to explore with systems thinking: an ex-post evaluation conducted in 2016, and commissioned by an international donor, the United States Agency for International Development (USAID). That ex-post evaluation confirmed that an education development intervention, the Kimberly Thusanang Programme (KTP) implemented between 1998 and 2006, resulted in sustained outcomes, which were directly linked to the KTP’s goal of improving school governance in the Francis Baard education district the Northern Cape.
The Master’s research builds on the ex-post evaluation’s analysis. Using qualitative data analysis, the researcher identified the types of sustainability found in the ex-post evaluation data set. Then, by applying Stockmann’s (1993a) ‘extended dynamic sustainability’ concept, the Master’s research found that the KTP intervention and some of its benefits were dynamically sustained through the general causal sustainability mechanisms of problem-solving, modelling and multiplication.
These findings are likely useful to research intervention sustainability, to design sustainable development interventions, and to evaluate intervention success. Further exploration of these general sustainability mechanisms needs to be conducted to determine if these mechanisms are generalisable to other development interventions and their sustained outcomes.

EvalEdge Podcast about Storytelling in Evaluation

 


EVALEDGE PODCATS LOGO

In this episode of EvalEdge, Asgar Bhikoo and I talk about Storytelling in Evaluation Practice. This episode focused on exploring current lessons related to the use of story-telling as innovation in evaluation practice in Africa. For more information, check out Digital Stories for Impact and Social Impact Storytelling: Using Impact Data to Drive Change.

Here is the impact story tool referenced in the podcast, also here on the Civicus repository



Monday, July 29, 2019

There are alternatives to Experimental and Quasi-Experimental Impact Evaluation Methods.


Some of my clients are really interested in measuring their impact. RCTs and other quasi-experiments are first on their list of suggested designs. But our repertoire of IE designs and methods have grown.




This DfID working paper says:


Most development interventions are ‘contributory causes’. They ‘work’ as part of a causal package in combination with other ‘helping factors’ such as stakeholder behaviour, related programmes and policies, institutional capacities, cultural factors or socio-economic trends. Designs and methods for IE need to be able to unpick these causal packages. 
Demonstrating that interventions cause development effects depends on theories and rules of causal inference that can support causal claims. Some of the most potentially useful approaches to causal inference are not generally known or applied in the evaluation of international development and aid. Multiple causality and configurations; and theory-based evaluation that can analyse causal mechanisms are particularly weak. There is greater understanding of counterfactual logics, the approach to causal inference that underpins experimental approaches to IE. 

Methods that I am currently interested in include 

Qualitative Impact Assessment Protocol 
The QuIP gathers evidence of a project’s impact through narrative causal statements collected directly from intended project beneficiaries. Respondents are asked to talk about the main changes in their lives over a pre-defined recall period and prompted to share what they perceive to be the main drivers of these changes, and to whom or what they attribute any change - which may well be from multiple sources.
Typically, a QuIP study involves 24 semi-structured interviews and four focus groups, conducted in the native language by highly-skilled, local researchers. However, this number is not fixed and will depend on the sampling approach used. The research team conducting interviews are independent and blindfolded where appropriate; they are not aware who has commissioned the research or which project is being assessed. This helps to mitigate and reduce pro-project and confirmation bias, as well as enable a broader and more open discussion with respondents about all outcomes and drivers of change.
Qualitative Comparative Analysis 
Qualitative Comparative Analysis (QCA) is a means of analysing the causal contribution of different conditions (e.g. aspects of an intervention and the wider context) to an outcome of interest. QCA starts with the documentation of the different configurations of conditions associated with each case of an observed outcome. These are then subject to a minimisation procedure that identifies the simplest set of conditions that can account all the observed outcomes, as well as their absence. The results are typically expressed in statements expressed in ordinary language or as Boolean algebra. QCA is able to use relatively small and simple data sets. There is no requirement to have enough cases to achieve statistical significance, although ideally there should be enough cases to potentially exhibit all the possible configurations. 

Friday, July 19, 2019

Picture this- Complexity

This handy poster made by Johanna Boehnert explains 16 terms that often pop up in thining about complex systems. It's a bit like a gateway drug to reading more on Complex Systems.

If found it in a tweet by @Heinomatti which refers to the website of CECAN .
But Better Evaluation also has a really nice summary of it.








Wednesday, July 17, 2019

Systems Science and Complexity Science - related but not the same

I'm studying again and for that, I'm reading. A lot. I'm reading about systems thinking and factors that support sustained outcomes of development interventions. Often I stumble on things that make me go: "Ooh - I should remember this next time I do ABC" So this blog is being revived a bit to help keep track of these random thoughts.

I read about the history of systems thinking and complexity science and how both fields have similar challenges. Two great resources:

Midgley and Richardson comparison of paradigms in the Systems Field and the Complexity Field. 



Midgley's reflection on the history of paradigm wars between systems scientists amongst themselves, and complexity scientists amongst themselves. He says: 

Systems scientists were embroiled in a paradigm war, which threatened to fragment the systems research community. This is relevant... because the same paradigms are evident in the complexity science community, and therefore it potentially faces the same risk of fragmentation.

My interest in reading about the relationship between systems science and complexity science got sparked when I looked for examples of emergence, feedback and self-organization in my data and couldn't figure out what that would look like. A colleague suggested that while the concept "feedback" definitely occurs in multiple branches of the systems field (oh and there are so very very many), that the concepts "emergence" and "self-organization" are from complexity science.

One may argue that it probably doesn't matter into which categories these concepts fall, but actually, it does. Because the ontological and epistemological assumptions that underly these paradigms may or may not be similar and should be questioned.

So to get my thinking about the concepts straight, I need to get my thinking about the paradigms straight. Its a work in progress....

Friday, January 27, 2017

Can you tell me "What works in..."

Although we reportedly now live in a post-evidence era, I still choose to cling to the minority view that programmes should be informed by research about what works. But where do you find the evidence?

About two years ago I attended a training course presented by Phil Davies from 3ie. He had many interesting insights to share, but today I was reminded of this excellent list of synthesised evidence that he shared.


One of my recent favourite systematic reviews, conducted by 3ie is The impact of education programmes on learning and school participation in low- and middle-income countries by Snilstveit et al . It has evidence about supplementary education programmes, feeding programmes, ICT in education programmes and a wide range of others. 


Happy Reading!  

Wednesday, July 13, 2016

Evaluative Rubrics - Helping you to make sense of your evaluation data

Three times in one week I've now found myself explaining the use of evaluation rubrics to potential evaluation users. I usually start with an example like this, that people can relate to:
When your high school creative writing paper was graded, your teacher most likely gave you an evaluative rubric which specified that you do well if you 1) used good grammar and spelling, 2) structured your arguments well, and 3) found an innovative and interesting angle on your topic. In essence, this rubric helped you to know what is "good" and what is "not good".
In an evaluation, a rubric does exactly the same. What is a good outcome if you judge a post- school science and maths bridging programme? How does the outcomes of "being employed" or  "busy with a third year  B Sc. Degree at university" compare to an outcome like "being a self-employed university drop-out with three registered patents" or to an outcome like "being unemployed and not sure what to do about the future". A rubric can help you to figure this out.

E. Jane Davidson has some excellent resources on rubrics here and here. If you need a rubric on evaluating value for investment, Julian King has a good resource here.  And of course, there is the usual great content on better evaluation here.

I love how Jane describes why we need evaluation rubrics:
Evaluative rubrics make transparent how quality and value are defined and applied. I sometimes refer to rubrics as the antidote to both ‘Rorschach inkblot’ (“You work it out”) and ‘divine judgment’ (“I looked upon it and saw that it was good”)-type evaluations.

Monday, February 29, 2016

Writing Summaries for Evaluation Reports


Last year I attended a course on "Using Evidence for Policy and Practice" presented by Philip Davies from the International Initiative for Impact Evaluation [3ie]. I found his guidelines for what should go into the 1:3:25 summaries most helpful. Here they are:
The full course material is available on the website of the African Evidence Network's Website. Here

Wednesday, October 07, 2015

What I'm up to at the 2015 SAMEA Conference

The SAMEA conference is happening from 12 to 16 October and I'm looking forward to it. 


5thSAMEA Conference LogoSince January, I've had to temporarily downscale my professional involvement in the M&E and Educational networks and I had to neglect this little blog a bit because of a second long term development project I took on in January 2015. The project has lovely brown eyes, an infectious laugh and goes by the name of Clarissa. I'm happy to report that no major clashes with the first development project, (Named Ruan) has so far occurred, but its been a bit of an adjustment to balance work, and volunteering, and life in general. 


So what am I up to  at the conference?
I'll be tweeting from @benitaw if you are interested in my perspective of the conference. I will also attend an IOCE stand at the conference, aiming to promote the VOPE Institutional Capacity Toolkit which my consultancy developed under the EvalPartners leadership of Jennifer Bisgard, Patricia Rogers, Jim Rugh, and Matt Galen. This is an online toolkit full of helpful resources aimed to equip VOPEs (Voluntary Organisations for Professional Evaluation) to become more accountable and more active. 

Then, I'll be teaming up with Cara Waller (from CLEAR) and Donna Podems (from OtherWise) in a session for African VOPEs  on Friday 16th October. This is a ‘world-café’ style event, from 10 –11:30am, to be held as a joint ‘Made in Africa’ and ‘Discussing the Professionalisation of Evaluation and Evaluators’ stream session.  The aim of the session is to provide a space for those involved with VOPEs in the region (and those with an interest in strengthening African VOPEs) to come together to discuss current topics around building quality supply and generating demand for evaluation in contextually-specific ways. So please come and chat all things VOPE on the day!

Good luck to my colleague Fazeela Hoosen and the rest of the SAMEA board on hosting this year's conference with the DPME and the PSC. I know (and boy.... do I know) it is very hard work. So thanks in advance for all of the hours you are putting in, to make this event happen. 

Thursday, October 16, 2014

True Confessions of an Economic Evaluation Phobic

You know how the forces at work in the universe sometimes conspire and confronts you with a persistent nudge... over an over again? Well this week's nudge was "You know nothing about economic evaluation... do something about it - Other than ignoring it".

Words like "cost-benefit analysis, cost-efficiency analysis, cost-utility analysis"... actually anything with the word "cost" or "expenditure" in it... makes me nervous. So my usual strategy is to ignore the "Efficiency" criterion suggested by the OECD DAC, or I start fidgeting around for the contact details of one of my economist friends, and pass the job along. I have even managed to be part of a team doing a Public Expenditure Tracking Survey without touching the "Expenditure" side of the data.

But then I found these two resources that helped me to start to make a little bit more sense of it all. They are:

The South African Department of Planning Monitoring and Evaluation's Guideline on Economic Evaluation  At least it starts to explain the very many different kinds of economic evaluation you should consider if you work within the context of South Africa's National Evaluation Policy Framework.


And then this. 
http://www.julianking.co.nz/downloads/


A free ebook by Julian King that presents a short theory for helping to answer the question "Does XYZ deliver good (enough) value for investment?" - Essentially the question any evaluator is supposed to help answer.

So, now, there is one more topic on my ever expanding reading list! If there is a "Bible" of economic evaluation, let me have the reference, ok?

Friday, September 12, 2014

What if, mid career as a researcher, you become interested in Evaluation?

An old classmate, that took the market research route after completing her Research Psych Master's Degree, asked me for a couple of references to check out if she wanted to develop her evaluation knowledge and skills. What came to mind is the following professional development resources. I'm sure there's many more easily accessible ones, but this is a good start for a list!