MEAL
Measuring what security work actually changes
Monitoring, evaluation, accountability and learning, applied to security and justice programmes. This page sets out what we do, what the evidence says about why it is hard, and what we will not claim to have proved.
Why we say MEAL and not MEL
The Foreign, Commonwealth and Development Office says MEL. Most of our clients and partners say MEL. We say MEAL deliberately, and the reason is not branding.
- In most sectors, accountability means reporting upwards to the donor.
- It is a duty owed by the implementer to the funder, discharged through quarterly returns. Useful, but it is administration.
- In our sector, accountability is the thing being built.
- The Geneva Centre for Security Sector Governance lists accountability as one of the seven principles of good security sector governance, alongside transparency, rule of law, participation, responsiveness, effectiveness and efficiency. A police force that cannot be held to account is not a police force that has been reformed.
- So the A is not a reporting obligation. It is an outcome.
- Does the complaints mechanism work. Does internal affairs sanction anyone. Does the parliamentary committee see the budget. Does the oversight body have powers it actually uses. Those are measurements, and they belong in the results framework rather than in an annex.
- We will use whichever term your framework uses.
- If your programme documentation says MEL, our reporting will say MEL. The distinction above is about what we measure, not about what we call it.
The four strands, as we use them
Monitoring
Continuous. Did the activity happen, to standard, on time, and what is the context doing around it. Monitoring answers whether the plan is being delivered. It does not answer whether the plan was right.
Evaluation
Periodic and structured. Did it work, why, and would we do it again. We work to the six OECD DAC criteria: relevance, coherence, effectiveness, efficiency, impact and sustainability.
Accountability
Two directions at once. Upwards to whoever is paying, in the format they require. And inside the institution, where the question is whether it has become more answerable to its own population.
Learning
Structured, not incidental. The development sector borrowed the after action review from military doctrine, and we use it as the army does: what was supposed to happen, what happened, why, and what changes tomorrow.
Why this is harder here than in health or education
Stated plainly, because a supplier who says measurement is straightforward in this sector has not done much of it.
- Success often looks like nothing happening.
- The clearest evidence is not a slogan but a finding. The independent impact evaluation of a large United Kingdom police reform programme in the Democratic Republic of the Congo found serious crime fell significantly while people’s sense of safety moved on a different timetable entirely. It warned that perceptions of change can run ahead of actual change, in both directions.
- Attribution is usually the wrong question.
- The field moved years ago from asking whether a programme caused an outcome to asking what it contributed to one. We design for contribution, test the causal pathway, and say so. Anyone promising to prove attribution in a conflict-affected environment is either inexperienced or selling.
- Counting outputs creates the wrong behaviour.
- Officers trained, courts built, vehicles delivered. A study of police performance measurement across several countries found that quantitative indicators can lead officers to decline to register cases that would damage their numbers. Aid reviewers have found the mirror image, where reach targets pushed providers to spread services too thinly to work.
- Capability scores go wrong in a specific and documented way.
- The United States special inspector general found that measurement methods for the Afghan security forces changed five times, making long-term tracking impossible, and that the highest rating available was "independent with advisors", which could not express the actual objective of a self-sustaining force. We cite that against ourselves, because scoring capability is what we do.
- The data is imperfect and the honest response is to say where.
- Records are incomplete, some regions are inaccessible, and the people producing the numbers sometimes have an interest in them. We grade our own confidence source by source and publish the grading.
What the record shows
Three findings that explain why we build assessment the way we do. None of them is ours.
The OECD Development Assistance Committee handbook section on monitoring and evaluating security system reform remains the most cited sector-specific guidance. It is nearly twenty years old.
The Independent Commission for Aid Impact scored United Kingdom security and justice assistance Amber-Red in 2015, and the Conflict, Stability and Security Fund Amber-Red in 2018. Measurement was central to both findings.
Community policing coverage fell from between 80 and 90 per cent in 2007 to 20 per cent by 2012, once external support ended. Reported by the same 2015 review.
How we build it into an engagement
Define the outcome before the activity
Not what will be delivered, but what will be different, for whom, and how anyone would know. Reviews of this sector repeatedly find programmes where outcomes were never defined, which makes later evaluation impossible rather than merely difficult.
Write the theory of change at the start, not halfway through
The Congo evaluation above records that the programme’s theory of change was written over halfway through implementation. That is common and it is costly, because it means the baseline was collected against assumptions nobody had yet written down.
Separate theory failure from implementation failure
A programme can be delivered exactly as designed and still not work, because the design was wrong. Telling those two apart is the single most useful thing an evaluation does, and it is only possible if the theory was written down in advance.
Set the baseline before delivery starts
A baseline collected after the first activity is not a baseline. Where a genuine baseline is impossible, we say so and design around it rather than reconstructing one later and hoping nobody asks.
Grade confidence, and publish the grading
Every finding carries a confidence level and a source. A reader should be able to discount anything we assert, and the weak points should be marked by us rather than found by them.
Build the human rights assessment into the evaluation, not beside it
The Overseas Security and Justice Assistance guidance requires a human rights risk assessment to be built into evaluation processes and updated when circumstances change. In practice it is often produced once, filed, and never revisited. Reviewers have found it approving activity to proceed unchanged in every case examined, which is a control that is not controlling anything.
Sustainment is a separate measurement problem
This is the part most frameworks leave out, and the part we think matters most.
- A strong result at the end of a programme does not predict a lasting one.
- The most rigorous study of post-exit sustainability found that impact at the point of exit does not consistently predict whether benefits survive, and that the size of the impact does not improve the odds either. That is not intuition. It is the finding that justifies measuring sustainment separately.
- Four things have to be in place before the money stops.
- Sustained resources, sustained capacity, sustained motivation, and linkages to something that outlasts the programme. In the same study, no project achieved sustainability without all of them in place before closure.
- We model the position at twelve and twenty-four months after support ends.
- Who pays the recurring costs, who holds the mandate, who is still in post, and what happens to the equipment when it needs its first major service. Salaries are usually the answer, and usually the thing nobody has funded.
- We will not pretend there is a body of evidence here.
- There is rigorous post-exit evidence in other sectors. For security and justice specifically there are strong individual cases and no systematic evidence base. We say that rather than implying otherwise, and it is a reason to measure rather than a reason not to.
What we will not claim
Working to your framework
Where a programme already has a results framework, a logframe or an evaluation contractor, we work inside it rather than proposing our own. Where there is none, we will build one proportionate to the size of the work, which for a small engagement means a short document rather than an apparatus. The Foreign, Commonwealth and Development Office judges evaluation against five principles, useful, credible, robust, proportionate, and safe and ethical, and proportionate is the one small suppliers most often fail in the wrong direction by over-engineering.