Guide
Baselines, endlines and evaluation designs
A baseline tells you where things started; an evaluation design tells you how much of the change was yours. Here are the designs, from simple to rigorous, when each is worth it, and what sample sizes really depend on.
At a glance11 min read
- Take the baseline before the work starts, and measure the endline the same way: same questions, same kind of sample and, where it matters, the same season.
- A before-and-after comparison credits a project with change that might have happened anyway. A comparison group shows what would have happened without it.
- Difference-in-differences, matching and cut-off designs can estimate a project’s effect without a lottery; randomised trials are the most rigorous but aren’t always possible or worth it.
- Qualitative and theory-based evaluation explains why and how change happened, and is often the right choice for small or new programmes.
- Sample size depends on how small a change you need to detect, how varied the outcome is, how clustered the sample is and how many people take part. Halving the change you want to detect needs four times the sample.
On this page
- Baselines, midlines and endlines
- The question every evaluation must answer
- Evaluation designs compared
- Before and after versus a comparison group
- Quasi-experimental methods in plain words
- Randomised trials: when they’re worth it
- Qualitative and theory-based evaluation
- Sample sizes in plain words
- Choosing an evaluator, budgets and ethics
- Questions people ask
- Sources
A baseline measures your key indicators before a project starts, and an endline measures them again at the end, so you can see what changed. An evaluation design is how you work out how much of that change the project caused, from a simple before-and-after comparison, through comparison groups and quasi-experimental methods, to randomised trials.
The right design depends on the question, the money and the moment. A small new programme needs good monitoring and an honest before-and-after picture. A large programme about to be scaled up, or one a company must assess under the CSR Rules, deserves a design that can separate its effect from everything else going on.
This guide is for NGO monitoring and evaluation staff, CSR teams commissioning studies and students. It explains baselines and endlines, the main designs in plain words, sample sizes, choosing an evaluator and the ethics of evaluation. For indicators, see choosing indicators; for the legal requirement, see impact assessment under the CSR rules.
Baselines, midlines and endlines
The OECD’s glossary defines a baseline as the conditions before an intervention, against which changes can be measured. In practice:
- The baseline measures your outcome indicators, and the characteristics you will break results down by, before activities begin.
- A midline, halfway through, checks whether the project is on track while there is time to change course.
- The endline measures the same indicators at the end.
- A follow-up, a year or more later, shows whether the change lasted. Under the CSR Rules, a mandatory impact assessment looks at projects completed at least a year earlier, so this is often when an assessor will visit.
The rule that matters most is consistency. Use the same questions, the same kind of sample, the same method and, where it matters, the same season. A nutrition survey in the lean season and another after the harvest will show change that has nothing to do with your project.
The question every evaluation must answer
Children’s reading improved. Would it have improved anyway? Teachers, the state’s learning programme, parents and other NGOs were all at work too. Evaluators call what would have happened without the project the counterfactual: the situation if there had been no intervention. You can never observe it directly, so every design is a way of estimating it.
This is also the difference between two kinds of claim. Attribution says how much of a change can be credited to the project. Contribution says, with evidence, that the project played a part alongside others. Most NGO and CSR projects can make a strong contribution claim; a firm attribution claim needs a design with a credible comparison.
Evaluation designs compared
| Design | How it works | When it’s credible | Main risk |
|---|---|---|---|
| Before and after | Measure the same group at the start and the end | Little else could have changed the outcome | Credits the project with change that would have happened anyway |
| Comparison group, at the end only | Compare participants with similar non-participants | The groups were alike to start with | They may have differed from the start |
| Difference-in-differences | Compare the change in participants with the change in a comparison group | Both groups were on similar trends before | Other things changed for one group only |
| Matching | Build a comparison group that resembles participants on measured traits | Rich data on who joined and why | Unmeasured differences, such as motivation |
| Cut-off (regression discontinuity) | Compare people just above and just below an eligibility score | The programme uses a clear cut-off | Says little about people far from the cut-off |
| Randomised trial | Decide by lottery who gets the programme first | A lottery is feasible and fair | Cost, time and spillovers between groups |
| Theory-based and qualitative | Test each link in the theory of change with mixed evidence | A comparison isn’t possible, or you need to know why | Can’t put a precise number on the effect |
Before and after versus a comparison group
Ten points is still a good result, and a far more believable one. A company that reported 19 points would have claimed nearly twice its real effect, and an assessor or a sharp-eyed CSR committee would ask why.
Quasi-experimental methods in plain words
Quasi-experimental designs estimate the counterfactual without a lottery. The four you will meet most:
- Difference-in-differences, as in the example. Its key assumption is that both groups would have changed at the same rate without the project. Check it: if you have data for earlier years, did the two groups move in parallel before?
- Matching finds non-participants who look like participants on things you can measure, such as age, caste, land or schooling. It can’t match on what you can’t see, such as a family’s determination, which is often why people join a programme in the first place.
- Cut-off designs use a rule. If scholarships go to students who score above 60, students who scored 59 and 61 are nearly identical, so comparing them shows the scholarship’s effect for students near that line.
- Phased roll-out. If the programme can’t reach every village at once, the villages that join next year are a natural comparison group this year, as long as the order wasn’t chosen because some villages were better prepared.
Randomised trials: when they’re worth it
In a randomised evaluation, also called a randomised controlled trial, participants or villages are assigned to the programme or a comparison group by chance, such as a lottery. J-PAL (the Abdul Latif Jameel Poverty Action Lab), a research centre known for this kind of study, explains why it works: random assignment makes the groups comparable, so differences at the end can be credited to the programme.
A trial is worth considering when:
- a big decision depends on the answer, such as scaling up or a government adopting the model;
- more people want the programme than it can serve, so a lottery is a fair way to choose;
- the programme is rolling out in phases anyway;
- the sample is large enough to detect the change you care about.
It isn’t the right tool when the programme is still changing, when the sample is too small, when the evidence wouldn’t change any decision, or when it would mean denying people something they are entitled to. Designs can often avoid that: a phased roll-out gives everyone the programme eventually, and an encouragement design randomly offers extra help to enrol without stopping anyone else from joining.
Qualitative and theory-based evaluation
Numbers tell you how much changed; people tell you why. Interviews, group discussions, observation and case studies explain what worked, for whom and in what circumstances, and they catch effects nobody planned, good or bad.
Theory-based approaches go further. Contribution analysis, for example, takes the project’s theory of change and tests each link with whatever evidence is available, to judge whether and how the project contributed to a result. For small or new programmes, where a comparison group isn’t possible, a careful theory-based evaluation is often the most useful and honest choice.
Qualitative work needs rigour too: choose whom to interview for a reason and say what it was, record and analyse systematically, look for evidence that contradicts your conclusions, and check findings against other sources. This cross-checking is called triangulation.
Sample sizes in plain words
How many households or children you need depends on what you want to know.
Detecting a change between two groups, as an impact evaluation must, needs careful power calculations. J-PAL’s guide to them makes four points worth remembering:
- Smaller effects need much bigger samples. Halving the effect you want to detect quadruples the sample you need.
- Clustering matters. If you assign whole schools or villages, the number of schools or villages matters a great deal, often more than the number of children in each.
- Take-up matters. If only a quarter of people offered the programme take it up, you need to offer it to 16 times as many people to detect the same effect as with full take-up.
- People drop out. Plan for families who move or can’t be found at the endline.
By convention, studies aim for an 80% chance of detecting a real effect, at a 5% significance level. If your budget can’t reach the sample this implies, say so, and choose a design that answers a question you can afford.
Choosing an evaluator, budgets and ethics
Independence. For a credible evaluation, the evaluator shouldn’t be the team that ran the project or anyone paid on its results. Agree at the start that findings will be reported as they are, good or bad. The detailed terms of reference for a mandatory CSR assessment are in our guide to impact assessment under the CSR rules.
Skills. Look for a field team that speaks the local language, statistical skills for the design you need, qualitative skills, sound ethics and data protection practice, and examples of past reports, including some with disappointing findings.
Budget. The cost is driven by the sample size, the number of survey rounds, distances and the season, any tests or measurements, and the time for analysis and writing. Each extra round repeats most of the field cost, so plan the baseline and endline together from the start. Under the CSR Rules, a company can count the cost of an impact assessment as CSR spending up to 2% of the year’s total CSR expenditure or ₹50 lakh, whichever is higher.
Ethics.
- Ask for informed consent, in the local language, and make clear that refusing won’t affect anyone’s access to services.
- For children, ask a parent or guardian and the child, and follow your safeguarding policy.
- Don’t deny a comparison group anything it would otherwise get; where you can, offer it the programme later.
- Protect personal data under the Digital Personal Data Protection Act, 2023 and its 2025 Rules, and remove names before sharing data with funders.
- For health measurements, such as children’s weight or anaemia tests, use trained staff, refer anyone who needs care, and seek ethics committee review.
- Share the findings with the community, not just the funder.
Questions people ask
- What is a baseline survey?
It is a measurement of a project’s key indicators before the work starts, such as the share of children reading at grade level or households with safe water. It is the starting point against which change is measured. Without one, a project can show what it did, but not what changed.
- What is the difference between a baseline and an endline?
The baseline is measured before the project begins and the endline at its end, using the same indicators, questions and methods. Comparing the two shows how much changed. To know how much of that change the project caused, you also need a comparison group or another way of estimating what would have happened anyway.
- What is a counterfactual in evaluation?
It is what would have happened to the same people or places without the project. It can’t be observed directly, so evaluations estimate it, for example with a comparison group of similar villages or schools. The project’s effect is the difference between what happened and the counterfactual.
- Do we need a randomised controlled trial?
Usually not. Randomised trials are most useful when a big decision depends on the answer, the programme is stable, a lottery is fair and the sample is large enough. Many programmes are better served by a comparison group design, a phased roll-out or a careful theory-based evaluation.
- How big should a baseline sample be?
To estimate a share to within 5 percentage points with 95% confidence, a simple random sample needs about 385 households. Sampling villages first means you need more, depending on how alike households in the same village are, and detecting a change between groups needs more again. Ask a statistician to calculate it for your design.
- What if we didn’t do a baseline?
Rebuild what you can from records made at the time, such as school registers, health records and enrolment forms, and use official data for the area. You can also measure a comparison group now and ask carefully worded recall questions. Be honest about the limits in your report, and take a proper baseline for the next phase.
Sources
- Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development (second edition, 2023) · OECD
- Introduction to randomized evaluations · J-PAL
- Quick guide to power calculations · J-PAL
- Companies (CSR Policy) Amendment Rules, 2022: impact assessment cost limit (First Notes, 19 October 2022) · KPMG in India
Go deeper in the Academy
Free course with a certificate · Intermediate · about 3½ hours
Is your CSR work worth a story?
The CSR Desk interviews CSR heads, and our weekly mailer reaches more than 50,000 readers.
Be seen for your workRead next
Impact and reporting
Impact assessment under the CSR rules
Larger companies must have an independent agency assess the impact of their bigger CSR projects. Here is exactly who is covered, the cost limit, the Social Stock Exchange exemption, and how to write terms of reference that produce a useful report.
Guide · 10 min read
Impact and reporting
Choosing indicators that mean something
An indicator is what you measure to know whether change is happening. Choose well and your reports show real results; choose badly and you end up counting kits and saplings. Here is how to choose, with examples by sector.
Guide · 10 min read
Impact and reporting
Needs assessment: start with the community
Good programmes start from what a community needs, not from what a funder wants to build. Here is how to find out, using India’s public data, honest conversations and local government, and how to turn what you learn into priorities.
Guide · 12 min read
Impact and reporting
Social return on investment (SROI) and its limits
SROI puts a rupee value on the changes a programme creates and compares it with what was spent. It can sharpen thinking about outcomes, but the headline ratio rests on many judgements. Here is how it works, and how to read one sceptically.
Explainer · 10 min read