Showing posts with label About Evaluation. Show all posts
Showing posts with label About Evaluation. Show all posts

Friday, July 03, 2026

What Guba and Lincoln offer to the use of AI debate

“Thou shalt use AI responsibly”





The current debate about AI often gets stuck in a fairly familiar place.

You can use AI, but you need to use it “well”.
Do not pass AI-generated work off as your own.
Do not let it do your analysis for you.
Do not let it write your report for you.
Be careful: it hallucinates references.

All of this is true. But I think it misses something more subtle.

AI, chocolate and cognitive offloading

For many evaluators, researchers and knowledge workers, the real risk is not that we deliberately ask AI to do work we know we should be doing ourselves. The real risk is that, when we are tired, under pressure, cognitively overloaded or facing a difficult piece of judgement, we slowly let the tool carry more of the thinking than we intended.

The issue is cognitive offloading.

Cognitive offloading is not new. Humans have always used external tools to reduce mental effort: notebooks, calculators, checklists, templates, search engines, colleagues, diagrams and frameworks. In that sense, AI is part of a long history of tools that extend our thinking.

But generative AI changes the nature of the offloading. It does not only store information or speed up a calculation. It can summarise, classify, infer, draft, compare, frame and argue. In other words, it can begin to look like analysis.

And here is the uncomfortable part: even when we know the risks, we are still human.

For anyone who has ever tried to diet while sitting next to chocolate, this is not hard to understand. You may have a clear rule. You may genuinely believe in the rule. You may even have defended the rule publicly. But when you are tired, stressed or depleted, the chocolate becomes much harder to resist.

AI is a bit like that.

The concern is not only: Did you use AI?

The more important question is: At what point did the judgement move?

This is where I think Guba and Lincoln have something surprisingly useful to offer.

 

Guba and Lincoln enters the room….

In a meeting recently, someone made a throwaway comment about Guba and Lincoln. I responded with a small verbal protest. This was probably lost on anyone who has wisely avoided the old evaluation “paradigm wars” — including the Lincoln–Sechrest debate, where Lincoln stood for constructivist, stakeholder-shaped evaluation, and Sechrest defended disciplined evidence, validity and causal reasoning.

Niche evaluation lore? Definitely. 

Still useful for thinking about AI? Surprisingly, yes.

But afterwards I kept thinking about it. Guba and Lincoln’s Fourth Generation Evaluation challenged the idea that evaluation quality should be judged only by the conventional research criteria of validity, reliability, generalisability and objectivity. For constructivist, qualitative and naturalistic inquiry, they proposed a different language of quality: credibility, transferability, dependability and confirmability.

These criteria have since become familiar shorthand for trustworthiness in qualitative research. And trustworthiness, it seems to me, is exactly where the AI debate needs to go.


 From “Thou shalt not” to practical questions

Telling people not to offload their thinking is probably not enough.

It is a bit like saying:

Do not be biased.
Do not make weak inferences.
Do not overclaim.
Do not eat the chocolate.

The instruction is correct, but not very operational.

Guba and Lincoln help translate the concern into practical checks. Let me start with one of their criteria: dependability.

Dependability: is the process traceable?

The dependability criterion asks whether the inquiry process is coherent, logical and traceable.

In qualitative research, this usually means that the process should be documented well enough for someone else to understand how the findings were produced. It does not require mechanical replication, but it does require an audit trail: records of key decisions, changes in focus, coding choices, analytic memos, and the movement from raw data to interpretation.

For AI use, I find dependability especially useful because it helps me ask:

Is the human analytic process still visible, or has the route from evidence to claim become blurred?

Application 1: using AI to strengthen a coding framework

When I use AI in qualitative analysis, I do not start by asking it to “find the themes”. I begin by reading through the data myself, making notes on possible themes, codes, tensions and patterns.

Only then do I use AI as a thinking partner. For example, I might ask:

“Let’s develop a deductive coding framework. Based on the themes I have identified, suggest a coding memo for each theme, possible sub-themes, and five quotes I can use to test whether the code is working.”

At that point, the work becomes iterative. I might respond:

“I like the distinction you are making here, but I think these two codes overlap.”
“This code is too broad.”
“This seems to miss what participants are saying about implementation conditions.”
“Let’s test this theme against more quotes.”

I may then move into a more inductive round:

“I want to see what the data says about this theme. Pull together all quotes that seem related to it, including quotes that complicate or contradict it.”

This process does not make the coding framework dependable because AI produced it. It becomes more dependable because the framework has been tested, revised and documented through several rounds of analytic questioning.

AI helps speed up the iteration, but the analytic judgement remains mine.

Application 2: checking whether an AI-supported executive summary is dependable

I also use the dependability criterion when AI helps me draft or refine an executive summary.

An executive summary can become polished very quickly. That is useful, but also risky. The more fluent the summary becomes, the easier it is to lose sight of where each claim came from.

So I ask AI to help me rebuild the audit trail:

“Please create a table listing each claim in the executive summary. For each claim, indicate which document, section, dataset or finding it is based on.”

I then check the table against the source material. I correct weak links, remove claims that are not well supported, and adjust wording where the claim has become too strong.

I may then ask:

“Based on the information I provided, are there any findings, stakeholder views or pieces of evidence that might contradict or complicate these claims?”

This helps me check whether the summary has become too neat. If there are counterclaims or important qualifications, I decide whether the executive summary needs to acknowledge them, soften the wording, or adjust the argument.

Only after that do I use AI for board-ready refinement:

“Please simplify this into a clearer board-ready message.”
“Make the tone more direct.”
“Reduce repetition.”
“Sharpen the main implication.”

Again, the key issue is not whether AI was used. The key issue is whether I can still trace the movement from evidence, to claim, to final wording.

For me, that is what dependability looks like in an AI workflow: not avoiding AI, but using it in a way that keeps the analytic path visible.

And then, where it matters, I put the evidence trail into annexes.

Often many, many pages of annexes.

Will everyone read them? Almost certainly not. But that is not the only point. The annexes make the reasoning visible and verifiable. They show how claims were built, what evidence they rest on, and where the qualifications sit.

They also create a written record that future machine synthesis will probably love: structured, source-linked, explicit reasoning that can be checked, reused and interrogated.

So yes, the annexes may be unloved by human readers.

But they are part of the dependability infrastructure.

Beyond dependability

I have focused here on dependability because it is the criterion that most directly helps me manage AI use in my workflow. It asks whether the process is traceable — whether the path from evidence, to interpretation, to final wording is still visible.

But the same logic applies to the other trustworthiness criteria.

For credibility, I use AI to test whether claims are well supported, whether counter-evidence has been considered, and whether the final account is believable in relation to the data.

For transferability, I check whether the context has survived the synthesis: whether the specific conditions, boundaries and setting are still visible, rather than being smoothed into a generic lesson.

For confirmability, I ask whether I can still distinguish between what came from the evidence, what came from my judgement, and what AI helped me formulate.

The bottom line

So for me, Guba and Lincoln do not provide a nostalgic evaluation reference. They offer a practical way of asking a very current question:

Not “Did you use AI?” but “Was the use of AI trustworthy?”**

 And if the answer is yes, can you show how?

**************************************************************

**Yes, Charlie and I wrote this blogpost together - the giveaway ?

The "Not this, but that" structure (also known as contrastive negation or antithetical parallelism) is a highly recognizable rhetorical pattern heavily favored by ChatGPT and other LLMs.

*** Photo by June Gathercole on Unsplash



Thursday, June 02, 2011

Evaluation Tasks

I found this graphical representation of Evaluation Tasks from Better Evaluation very useful for thinking about the evaluation process.


(Click on the pic for a larger version).


In my experience the "synthesize findings across evaluations"-bit gets neglected. In my work as an evaluator contracted to many corporate donors, I am usually required to submit an evaluation report for use by the client. I often have to sign a confidentiality agreement that prohibits me from doing any formal synthesis and sharing, even if I am doing similar work for different clients. Informally, I do share from my experience, but the communication is based on my anecdotal retellings of evidence that has been integrated in a very patchy manner.I try to push and prod clients into talking to each other about common issues, but this rarely results in a formal synthesis.
  
It is not always feasible for the clients who commission evaluations to do this kind of synthesis. Their in-house evaluation capacity rarely includes the meta-analysis skill, and even if they contract a consultant to conduct a meta-analysis based on a variety of their own evaluations, there are some problems: Aggregating findings from a range of evaluations that do not pay attention to the possibility that a meta-analysis will be done somewhere in the future, requires a bit of a “fruit-salad approach” where apples and oranges, and even some peas and radishes, are thrown together. Another obvious problem is that donors who do not care to share the good, bad and ugly of their programs with the entire world, would be hesitant to make their evaluations available for a meta-analysis conducted by another donor. 

Perhaps we require a “harmonization” effort among the corporate donors working in the same area?

Wednesday, February 06, 2008

Taxonomy of Evaluation

I found the Evaluation Webring's Taxonomy of Types, Approaches and Fields of Evaluation Quite Useful.

Category 1: Types of evaluation
Internal evaluation or self- evaluation
An evaluation carried out by members of the organisation(s) who are associated with the programme, intervention or activity to be evaluated.

Ex-ante evaluation or impact assessment
An assessment which seeks to predict the likelihood of achieving the intended results of a programme or intervention or to forecast its unintended effects. This is conducted before the programme or intervention is formally adopted or started. Common examples of ex-ante evaluation are environmental and/or social impact assessments and feasibility studies.

Mid-term or interim evaluation
An evaluation conducted half-way through the lifecycle of the programme or intervention to be evaluated. Monitoring An ongoing activity aimed at assessing whether the programme or intervention is implemented in a way that is consistent with its design and plan and is achieving its intended results.

Ex-post or summative evaluation
An evaluation which usually is conducted some time after the programme or intervention has been completed or fully implemented. Generally its purpose is to study how well the intervention served its aims, and to draw lessons for similar
interventions in the future.

Meta-evaluation
Two processes are often referred to as meta-evaluation: (1) the assessment by a third evaluator of evaluation reports prepared by other evaluators; and (2) the assessment of the performance of systems and processes of evaluation.

Formative evaluation
An evaluation which is designed to provide some early insights into a programme or intervention to inform management and staff about the components that are working and those that need to be changed in order to achieve the intended objectives.

Category 2: Evaluative approaches
Outcome evaluation
An evaluation which is focused on the change brought about by the programme or intervention to be evaluated or its results regarding the intended beneficiaries.
Impact evaluation An evaluation that focuses on the broad, longer-term impact or effects, whether intended or unintended, of a programme or intervention. It is usually done some time after the programme or intervention has been completed.

Performance evaluation
An analysis undertaken at a given point in time to compare actual performance with that planned in terms of both resource utilization and achievement of objectives. This is generally used to redirect efforts and resources and to redesign structures.

Participatory evaluation
An evaluation that actively involves all or selected stakeholders in the evaluation process. Different approaches involve varying degrees of participation, inclusion, capacity-building, ownership, etc.

Empowerment evaluation
An approach that aims to improve programs through using specific tools for assessing the planning, implementation and self-evaluation of programs, and by incorporating evaluation into a program or organization’s planning and management. It involves a high level of participation by stakeholders in the evaluation process and is guided by ten key principles.
Collaborative evaluation An evaluation which aims for a significant degree of collaboration or cooperation between evaluators and stakeholders.

Utilization-focused evaluation
A process that assists the primary intended users of an evaluation to select the most appropriate content, model, methods, and theory for the evaluation, focusing on their intended use of the evaluation. Use refers to how people apply evaluation findings and experience the evaluation process.

Feminist evaluation
An evaluation that commonly involves adapting or redesigning relevant evaluation theories and methodologies so that they are compatible with feminist theories and methodologies. Feminist evaluations aim to be inclusive and empowering for women in particular.

Theory-based evaluation
An evaluation based on the theories of change that underlie a given programme or intervention. Its major aim is to examine the extent to which these theories hold and to validate their underlying assumptions.

Most Significant Change
A form of participatory monitoring and evaluation which involves the collection and systematic review and analysis of change stories by panels of designated stakeholders or staff. It is mainly used to assess intermediate program impacts and outcomes.

Category 3) Fields of evaluation

Programme or project evaluation
The evaluation of a programme or project.

Policy evaluation
The evaluation of policies and procedures.

Evaluation of legislation
The evaluation of a piece of legislation.

Evaluation of technical assistance
The evaluation of technical assistance provided by international, bilateral or multilateral donors.

Organisation or institutional evaluation
An evaluation of an organization’s or other institution’s capacity for innovation and change. It involves examining its decision-making processes and organisational structures.

Proposal assessment
The assessment of bids presented by tenderers following a specific call for tenders/bids.
Financial audit The scrutiny of accounts of an organization or other institution against a set of standards.

Personnel evaluation
A systematic method of evaluating an employee’s or staff member’s performance. This involves tracking, evaluating and providing feedback in relation to specific predetermined standards which are consistent with the organization’s overall

Wednesday, January 09, 2008

competencies/capabilities of an evaluator

Q: Do you know of a document that articulates competencies/capabilities of an evaluator? If you do please send me a reference or copy.

A: Lots of work has been done on this topic by various Evaluation Associations across the world. Some useful references:

King, Jean, Stevahn, Laurie, Ghere, Gail, & Minnema, Jane (2001). Toward a taxonomy of essential evaluator competencies. American Journal of Evaluation, 22, 229-247.

Mertens, Donna M. (1994). Training evaluators: Unique skills and knowledge. New Directions for Program Evaluation, 62, 17-27.

Treasure Board of Canada Competency Profile for Federal Public Service Evaluation Professionalshttp://www.tbs-sct.gc.ca/eval/dev/Professionalism/profession_e.asp

Here is the list:

Essential Competencies for Program Evaluators (ECPE)

(Stevahn and King, Ghere, & Minnema, American Journal of Evaluation, March

2005)

1.0 Professional Practice

1.1 Applies professional evaluation standards

1.2 Acts ethically and strives for integrity and honesty in conducting evaluations

1.3 Conveys personal evaluation approaches and skills to potential clients

1.4 Respects clients, respondents, program participants, and other stakeholders

1.5 Considers the general and public welfare in evaluation practice

1.6 Contributes to the knowledge base of evaluation


2.0 Systematic Inquiry

2.1 Understands the knowledge base of evaluation (terms, concepts, theories, assumptions)

2.2 Knowledgeable about quantitative methods

2.3 Knowledgeable about qualitative methods

2.4 Knowledgeable about mixed methods

2.5 Conducts literature reviews

2.6 Specifies program theory

2.7 Frames evaluation questions

2.8 Develops evaluation designs

2.9 Identifies data sources

2.10 Collects data

2.11 Assesses validity of data

2.12 Assesses reliability of data

2.13 Analyzes data

2.14 Interprets data

2.15 Makes judgments

2.16 Develops recommendations

2.17 Provides rationales for decisions throughout the evaluation

2.18 Reports evaluation procedures and results

2.19 Notes strengths and limitations of the evaluation

2.20 Conducts meta-evaluations


3.0 Situational Analysis

3.1 Describes the program

3.2 Determines program evaluability

3.3 Identifies the interests of relevant stakeholders

3.4 Serves the information needs of intended users

3.5 Addresses conflicts

3.6 Examines the organizational context of the evaluation

3.7 Analyzes the political considerations relevant to the evaluation

3.8 Attends to issues of evaluation use

3.9 Attends to issues of organizational change

3.10 Respects the uniqueness of the evaluation site and client

3.11 Remains open to input from others

3.12 Modifies the study as needed


4.0 Project Management

4.1 Responds to requests for proposals

4.2 Negotiates with clients before the evaluation begins

4.3 Writes formal agreements

4.4 Communicates with clients throughout the evaluation process

4.5 Budgets an evaluation

4.6 Justifies cost given information needs

4.7 Identifies needed resources for evaluation, such as information, expertise, personnel, instruments

4.8 Uses appropriate technology

4.9 Supervises others involved in conducting the evaluation

4.10 Trains others involved in conducting the evaluation

4.11 Conducts the evaluation in a nondisruptive manner

4.12 Presents work in a timely manner


5.0 Reflective Practice

5.1 Aware of self as an evaluator (knowledge, skills, dispositions)

5.2 Reflects on personal evaluation practice (competencies and areas for growth)

5.3 Pursues professional development in evaluation

5.4 Pursues professional development in relevant content areas

5.5 Builds professional relationships to enhance evaluation practice


6.0 Interpersonal Competence

6.1 Uses written communication skills

6.2 Uses verbal/listening communication skills

6.3 Uses negotiation skills

6.4 Uses conflict resolution skills

6.5 Facilitates constructive interpersonal interaction (teamwork, group facilitation, processing)

6.6 Demonstrates cross-cultural competence

Thursday, August 16, 2007

What is an Evaluator?

Do you find it difficult to explain to people what you do? Those magical two sentences that will get people to go "Aaaaaah, now I get what you do?" Unfortunately I have not been able to come up with something concrete yet. But I am still trying.

In a television interview to publicise the SAMEA conference, the DDG from the PSC Mr. Mash Dipofu tried to explain it with an example. He asked the TV presenter if he knew what the viewers thought of his programme, how it can be improved and how many people actually watches it. He explained that by answering these questions, you are doing what an evaluator would be doing and answering the underlying question: Does what I am doing have value?

Which got me thinking. So much of what we do as evaluators are also done by other professionals.

  • We are a little like investigative journalists: We talk to people and ask questions and gather information to make an argument for or argainst something. Sometimes to inform readers of some wrong doing... Sometimes we celebrate what has been achieved.
  • Then we are also a little like the weather guy. We collect numbers over a long period of time and by applying some statistical techniques we can start predicting what will happen in future.
  • Another way of looking at our job is to compare it to that of teachers. We guide people to learn from their environments - assuming that they need to be taught how to use the information at their disposal to make intelligent choices. We check with tests whether the intended result has been achieved... much like teachers check whether their students have mastered a skill or knowledge component.
  • And then of course evaluators are also a little like an auditor in the way that we try to prove to people that money has been well spent.

The problem with explaining to people what we do, is probably because people tend to confuse it with research and planning and implementation and all kinds of other things. I have also spent some time to think about how being an evaluator is different from being a researcher, a planner and an implementer.

  • Because we use research techniques to collect evidence in order to evaluate, the difference between being a researcher (who asks questions in a specific way to gather evidence) and an evaluator (who asks questions in a specific way to gather evidence to then make a value judgment about the evaluand) is sometimes a little difficult to explain. But there is a difference!
  • Planning, on the other hand comes quite naturally when you are an evaluator. After delivering an evaluation, I frequently get asked to assist in planning processes - if people value what you produced in the evaluation they want to make sure that they plan to implement the recommendations made. Being a weather guy and an auditor makes it easier to plan because you can draw info together to make predictions, and you know that you will have to explain to people why you chose to spend their money in a particular way.
  • I think evaluators will probably make terrible implementers. As an evaluator you are constantly asking questions: Is this the best way to do things? Will we achieve results? How would we know that we added value? How do we know that this is the best way forward? To implement, however, you sometimes have to say "Well I don't know all the answers but I am making a decision to do ABC in the following way and that is the way it is!"
Despite having useful analogies to explain what evaluators do and don't do, I think that we are at risk if we, as practitioners of a scientific metadiscipline, don't understand how the ideology underlying evaluation is different from other those informing other jobs. Evaluators might be sharing some commonalities with teachers, weather people, auditors and journalists, but we have different values and assumptions guiding our work. We cannot forget that our work probably has a deeply political nature because we have to choose at some stage whose questions we will have to answer.

I found the following useful bit about evaluation approaches and the underlying philosophy, epistemology and ontology at http://www.recipeland.com/facts/Evaluation

Classification of approaches

Two classifications of evaluation approaches by House House, E. R. (1978). Assumptions underlying evaluation models. Educational Researcher. 7(3), 4-12. and Stufflebeam & Webster Stufflebeam, D. L., & Webster, W. J. (1980). An analysis of alternative approaches to evaluation. Educational Evaluation and Policy Analysis. 2(3), 5-19. can be combined into a manageable number of approaches in terms of their unique and important underlying principles.

House considers all major evaluation approaches to be based on a common ideology, liberal democracy. Important principles of this ideology include freedom of choice, the uniqueness of the individual, and empirical inquiry grounded in objectivity. He also contends they all are based on subjectivist ethics, in which ethical conduct is based on the subjective or intuitive experience of an individual or group. One form of subjectivist ethics is utilitarian, in which â€Å“the good” is determined by what maximizes some single, explicit interpretation of happiness for society as a whole. Another form of subjectivist ethics is intuitionist / pluralist, in which no single interpretation of â€Å“the good” is assumed and these interpretations need not be explicitly stated nor justified.

These ethical positions have corresponding epistemologies—philosophies of obtaining knowledge. The objectivist epistemology is associated with the utilitarian ethic. In general, it is used to acquire knowledge capable of external verification (intersubjective agreement) through publicly inspectable methods and data. The subjectivist epistemology is associated with the intuitionist/pluralist ethic. It is used to acquire new knowledge based on existing personal knowledge and experiences that are (explicit) or are not (tacit) available for public inspection.

House further divides each epistemological approach by two main political perspectives. Approaches can take an elite perspective, focusing on the interests of managers and professionals. They also can take a mass perspective, focusing on consumers and participatory approaches.

Stufflebeam and Webster place approaches into one of three groups according to their orientation toward the role of values, an ethical consideration. The political orientation promotes a positive or negative view of an object regardless of what its value actually might be. They call this pseudo-evaluation. The questions orientation includes approaches that might or might not provide answers specifically related to the value of an object. They call this quasi-evaluation. The values orientation includes approaches primarily intended to determine the value of some object. They call this true evaluation.

http://www.recipeland.com/facts/Evaluation



Monday, January 29, 2007

Making Evaluation our Own

At the AfrEA conference, there was a special stream on: 'Making Evaluation our Own'. It aimed to investigate where we are in terms of having Africa rooted, Africa lead evaluations.

I found it particularly useful because it became patently obvious that there are African world views and African methods of knowing that are not yet exploited for Evaluation in Africa. This of course brings the whole debate about "African" Evaluation theories to bear, and asks which kinds of evaluation theories are currently influencing our practice as evaluators in Africa.

Marvin C. Alkin and Christina A. Christie developed what they call the EVALUATION THEORY TREE. It splits the prominent (North-American) evaluation theorists into three big branches: Theories that focus on the use of evaluation, theories that focus on the methods of evaluation and theories that focus on how we value when evaluating. You can find more information about this at http://www.sagepub.com/upm-data/5074_Alkin_Chapter_2.pdf




The second tree is a slightly updated version. It was interesting to note that most of my reading about evaluation has been on "Methods" and "Use".

I think that if we are serious about developing our own African evaluation theories, we might need to develop our own African tree. Bob Piccioto mentioned that the African tree might use the branches of the above tree as roots, and grow its own unique branches.


A small commission from the conference put together a call for Action that outlines some key steps that should be taken if we hope to make progress soon. Hopefully I can post this at a later stage.

Keep well!

Thursday, January 25, 2007

UFE & The difference between Evaluation and Research

At the recent AFREA conference I was again reminded of what we are supposed to be doing in evaluation. Consider the word evaluation: It is about valuing something. Valuing for the purposes of accountability and for learning and improvement.

It is not just research, and although some people have indicated that they get irritated with our attempts at distinguishing evaluation from research, I think it is critically important to distinguish between research and evaluation.

Depending on which paradigm you come from, one might argue that research can be the same as evaluation. I don’t argue with that. What I do have a problem with is people approaching evaluations like research projects where the focus is all on “How do we collect evidence?” The methodology is critically important, agreed, and there is nothing that grates me more than seeing how people use poorly designed evaluation methodologies to collect “evidence”.

But evaluation is not just about how we collect information. Evaluation is supposed to take it a step further and make some evaluative judgments based on the data that was collected. Just describing your evaluation findings without saying what it means is senseless.

It is good and well if you find information about the level of maths capacity in rural schools interesting, but an evaluation will also go further and indicate whether the project is relevant, effective, efficient, has an impact and is sustainable or creates sustainable results. Without this additional “Valuing” judgments, an evaluation is only a research project that may increase our knowledge, but don’t help us to make decisions.

Something that may help more evaluations to be true evaluations is the Utilization Focused Evaluation approach of Michael Quinn Patton. It is all about how to ensure that an evaluation serves its intended purpose for the intended users. Go ahead – google Utilization Focused Evaluation and see how many hits come up. It literally is the biggest thing that has hit the Evaluation community in the past 30 years, yet many people are blissfully ignorant of this.

For those who commission evaluations, Patton specifically created a checklist that may be of value in making sure that evaluations are useful. www.wmich.edu/evalctr/checklists/ufe.pdf It might need to be adapted for use in your specific setting, but it definitely asks a couple of pretty critical questions about our evaluations.


Go ahead… I dare you to read up more about UFE (Utilization Focused Evaluation) and not be excited about the possibilities that evaluation has!


Have A good day!

PS. I hope to post some more of my thoughts on the AfrEA conference over the next month or so!

Thursday, January 04, 2007

IOCE

The IOCE is an international organisation for cooperation in evaluation and they have a couple of neat resources on their website:

http://www.ioce.net/resources/reports.shtml

The World Bank's Independent Evaluation Group Finds Progress On Growth, But Stronger Actions Needed For Sustainable Poverty Reduction



The World Bank's Independent Evaluation Group (IEG) is releasing its 2006 Annual Report on Operations Evaluation (AROE)

Joint UNICEF/IPEN Evaluation Working Paper on "New trends in development evaluation"

Resources for Evaluation and Social Research Methods

What Constitutes Credible Evidence in Evaluation and Applied Research? >>

When Will We Ever Learn: Recommendations to Improve Social Development through Enhanced Impact Evaluation

Monday, October 16, 2006

Common Pitfalls in M&E

This is an outline for a presentation I recently deliverd.

Common Pitfalls in Monitoring and Evaluation
Issues to Consider when you are the implementer / commissioner of evaluations

Introduction: What people Think of Evaluations
Often people are very scared of evaluations because of previous experiences, lack of experience or a general misconception regarding evaluations.

Introduction: Why must we measure?
Although there is growing consensus that we need to measure the results (outputs, outcomes and impacts) of our projects / programmes / policies, there is still much confusion about exactly why we are doing it.
Two main purposes of evaluations:
-- Accountability to various stakeholders
--Learning to improve the projects / programmes / policies
The projects / progammes / policies we implement affect thousands of people and if we get it wrong thousands will be affected negatively (or not affected at all)
We often complain about the cost of measuring our impact, but have we considered the costs of not measuring our impact?


Introduction: We want to evaluate BUT…
Once we are convinced that we should be measuring our impacts, a range of other questions come up:
--How should it be evaluated?
--When should it be evaluated?
--How will we know that the impact is the best possible?
--How do we know if it is our programme that made those differences?
--Can we do our own evaluation or should we get some specialist to do it?
--If there were simple one-size fits all answers to these questions, evaluation would probably have been much more appealing than it is today.

Common Pitfalls in Evaluation 1
Failing to clarify the intended use or the intended users of the evaluation – Producing "Door Stops".
Thinking you can evaluate your impact after year one of an intervention in a complex system – Expecting too much.
Thinking your impact evaluation is only something you need to worry about at the end of the project – Waiting too long.
Measuring every detail of a programme thinking that it will allow you to get to the big picture "impact" – Measuring too much.
Doing the wrong type of evaluation for the phase in which the project is in – Method / timing match.


Common Pitfalls in Evaluation 2
Allocating too little time and resources to the evaluation – More is better.
Allocating too much time and resources to the evaluation - Less is more.
Sticking to your or someone else’s "template" only – One size does not fit all.
Thinking that an online M&E system will solve all of your problems – Computers don’t solve everything.
Not planning for how the evaluation findings will be used – Findings don’t speak for themselves.

Common Pitfalls in Evaluation 3
Running a lottery when you are supposed to receive tenders for doing the evaluation – Lottery evaluations
Sending the evaluation team in to open Pandora’s box – Don’t do evaluation if you need Organisational Development.
Doing an impact evaluation without taking into consideration the possible influence of other initiatives / factors in the environment – Attribution Error.
Doing an impact evaluation without looking what the unintended consequences of the project was – Tunnel Vision
Ignoring the voices of the "evaluated" – Disempowering people

Common Pitfalls in Evaluation 4
Expecting your content specialist to also be an evaluation specialist and vice-versa – Pseudo Specialists lead to pseudo knowledge
Doing evaluations, creating expectations and then ignoring the results
Do not report statistics like level of significance and effect size when you incorporate a quantitative aspect to your evaluation – Being afraid of the "hard stuff"
Do not acknowledge the lenses you are using to analyse your qualitative data – Being colour blind
Getting hung up on the debate about whether quantitative / qualitative methods are better – Method Madness

How to address the pitfalls
Given that until very recently there were no academic programmes focusing on training people in evaluation, it is important that we find ways of improving our understanding of the field.
You need not be an evaluation specialist to be involved with evaluation.
Make sure that the evaluators you work with have development as an ultimate goal.


How to address the pitfalls
Resources for helping you to do / commission better evaluations
Join an association: For example the South African Monitoring and Evaluation Association (http://www.samea.org.za/) or the African Evaluation Association (http://www.afrea.org/)
Take cognisance of the guidelines and standards produced by these organisations
Make use of the many online resources available on the topic of evaluation (Check out Resources on the SAMEA web page)

Thursday, August 17, 2006

Cultural Competence of Evaluators

Hazel Symonette from the University of Wisconsin recently visited South Africa and presented M&E workshops in collaboration with the South African Monitoring and Evaluation Association. Unfortunately my diary did not allow me to attend any of the workshops, but I was lucky enough to have some interaction with her on an informal basis. This made me think about cultural competence required by evaluators. Look, we are long past the positivistic view where an evaluator was believed to be the expert able to look at behaviour and responses of people and categorise it objectively. What Hazel’s visit reinforced for me was the fact that cultural competence and identifying the lenses through which we look is extremely important if we want to do a good job as an evaluator.

This morning I read an article in the paper about learners in Mpumalanga schools:

'Teachers are bewitching us' 2006-08-16 19:07:56 http://www.mweb.co.za/news/?p=top_article&i=224129

There appears to be a growing tendency among Mpumalanga school pupils to accuse their teachers of witchcraft and then start a riot or boycott class. Nelspruit - There's a growing tendency among Mpumalanga school pupils to accuse their teachers of witchcraft and then start a riot or boycott class. Pupils at four schools have rioted in separate incidents since March, said provincial education spokesperson Hlahla Ngwenya on Wednesday. The latest incident happened on Monday when pupils at Mambane secondary school in Nkomazi, south of Malelane, refused to attend classes after allegations that teachers were bewitching them. The pupils returned to class on Tuesday. "Our preliminary reports indicate that the pupils protested after some of their peers died in succession over a short period," said Ngwenya. "They seem to believe this was the doing of their teachers." He said the department was investigating the incident and that pupils found guilty of instigating the boycott faced expulsion.


Imagine I was an evaluator in that community, working with the schools on the evaluation of some whole school development initiative. From my Westernised perspective witchcraft is just silly, and people believing in witchcraft are obviously mistaking one issue for another. Do I have the competence to be the evaluator in such a situation? How valid would my conclusions have been if I was in that situation?

I would probably have searched for alternative explanations, or more culturally acceptable explanations – I.e. There is obviously a problem in the relationship between the educators and the learners. It also seems that there are a range of very unfortunate circumstances (possibly a problem with HIV/AIDS?) in that community that needs attention. Just because I don’t accept their explanation and choose to come up with other explanations that are more culturally acceptable in my frame of reference (and probably in the frame of reference from which the programme donors come), does that mean it is the correct answer? Isn’t there maybe something beyond my perspective?

In my time as an evaluator I have come across a couple of other similarly absurd sets of behaviours – Teachers that toyi-toyi about catering whilst being on a government sponsored training session. Project beneficiaries refusing to disclose their names during interviews about an NGO’s performance. Clients being scared of saying anything out of fear that they might experience negative circumstances. Maybe these “absurdities”, when I recognize them, is a cue that I am out of my league?