Skip to content

Guide

Baselines, endlines and evaluation designs

A baseline tells you where things started; an evaluation design tells you how much of the change was yours. Here are the designs, from simple to rigorous, when each is worth it, and what sample sizes really depend on.

SocioStory Knowledge desk

Reviewed 11 min read

At a glance11 min read

  • Take the baseline before the work starts, and measure the endline the same way: same questions, same kind of sample and, where it matters, the same season.
  • A before-and-after comparison credits a project with change that might have happened anyway. A comparison group shows what would have happened without it.
  • Difference-in-differences, matching and cut-off designs can estimate a project’s effect without a lottery; randomised trials are the most rigorous but aren’t always possible or worth it.
  • Qualitative and theory-based evaluation explains why and how change happened, and is often the right choice for small or new programmes.
  • Sample size depends on how small a change you need to detect, how varied the outcome is, how clustered the sample is and how many people take part. Halving the change you want to detect needs four times the sample.
On this page
  1. Baselines, midlines and endlines
  2. The question every evaluation must answer
  3. Evaluation designs compared
  4. Before and after versus a comparison group
  5. Quasi-experimental methods in plain words
  6. Randomised trials: when they’re worth it
  7. Qualitative and theory-based evaluation
  8. Sample sizes in plain words
  9. Choosing an evaluator, budgets and ethics
  10. Questions people ask
  11. Sources

A baseline measures your key indicators before a project starts, and an endline measures them again at the end, so you can see what changed. An evaluation design is how you work out how much of that change the project caused, from a simple before-and-after comparison, through comparison groups and quasi-experimental methods, to randomised trials.

The right design depends on the question, the money and the moment. A small new programme needs good monitoring and an honest before-and-after picture. A large programme about to be scaled up, or one a company must assess under the CSR Rules, deserves a design that can separate its effect from everything else going on.

This guide is for NGO monitoring and evaluation staff, CSR teams commissioning studies and students. It explains baselines and endlines, the main designs in plain words, sample sizes, choosing an evaluator and the ethics of evaluation. For indicators, see choosing indicators; for the legal requirement, see impact assessment under the CSR rules.

Baselines, midlines and endlines

The OECD’s glossary defines a baseline as the conditions before an intervention, against which changes can be measured. In practice:

  • The baseline measures your outcome indicators, and the characteristics you will break results down by, before activities begin.
  • A midline, halfway through, checks whether the project is on track while there is time to change course.
  • The endline measures the same indicators at the end.
  • A follow-up, a year or more later, shows whether the change lasted. Under the CSR Rules, a mandatory impact assessment looks at projects completed at least a year earlier, so this is often when an assessor will visit.

The rule that matters most is consistency. Use the same questions, the same kind of sample, the same method and, where it matters, the same season. A nutrition survey in the lean season and another after the harvest will show change that has nothing to do with your project.

The question every evaluation must answer

Children’s reading improved. Would it have improved anyway? Teachers, the state’s learning programme, parents and other NGOs were all at work too. Evaluators call what would have happened without the project the counterfactual: the situation if there had been no intervention. You can never observe it directly, so every design is a way of estimating it.

This is also the difference between two kinds of claim. Attribution says how much of a change can be credited to the project. Contribution says, with evidence, that the project played a part alongside others. Most NGO and CSR projects can make a strong contribution claim; a firm attribution claim needs a design with a credible comparison.

Evaluation designs compared

DesignHow it worksWhen it’s credibleMain risk
Before and afterMeasure the same group at the start and the endLittle else could have changed the outcomeCredits the project with change that would have happened anyway
Comparison group, at the end onlyCompare participants with similar non-participantsThe groups were alike to start withThey may have differed from the start
Difference-in-differencesCompare the change in participants with the change in a comparison groupBoth groups were on similar trends beforeOther things changed for one group only
MatchingBuild a comparison group that resembles participants on measured traitsRich data on who joined and whyUnmeasured differences, such as motivation
Cut-off (regression discontinuity)Compare people just above and just below an eligibility scoreThe programme uses a clear cut-offSays little about people far from the cut-off
Randomised trialDecide by lottery who gets the programme firstA lottery is feasible and fairCost, time and spillovers between groups
Theory-based and qualitativeTest each link in the theory of change with mixed evidenceA comparison isn’t possible, or you need to know whyCan’t put a precise number on the effect

Before and after versus a comparison group

Ten points is still a good result, and a far more believable one. A company that reported 19 points would have claimed nearly twice its real effect, and an assessor or a sharp-eyed CSR committee would ask why.

Quasi-experimental methods in plain words

Quasi-experimental designs estimate the counterfactual without a lottery. The four you will meet most:

  • Difference-in-differences, as in the example. Its key assumption is that both groups would have changed at the same rate without the project. Check it: if you have data for earlier years, did the two groups move in parallel before?
  • Matching finds non-participants who look like participants on things you can measure, such as age, caste, land or schooling. It can’t match on what you can’t see, such as a family’s determination, which is often why people join a programme in the first place.
  • Cut-off designs use a rule. If scholarships go to students who score above 60, students who scored 59 and 61 are nearly identical, so comparing them shows the scholarship’s effect for students near that line.
  • Phased roll-out. If the programme can’t reach every village at once, the villages that join next year are a natural comparison group this year, as long as the order wasn’t chosen because some villages were better prepared.

Randomised trials: when they’re worth it

In a randomised evaluation, also called a randomised controlled trial, participants or villages are assigned to the programme or a comparison group by chance, such as a lottery. J-PAL (the Abdul Latif Jameel Poverty Action Lab), a research centre known for this kind of study, explains why it works: random assignment makes the groups comparable, so differences at the end can be credited to the programme.

A trial is worth considering when:

  • a big decision depends on the answer, such as scaling up or a government adopting the model;
  • more people want the programme than it can serve, so a lottery is a fair way to choose;
  • the programme is rolling out in phases anyway;
  • the sample is large enough to detect the change you care about.

It isn’t the right tool when the programme is still changing, when the sample is too small, when the evidence wouldn’t change any decision, or when it would mean denying people something they are entitled to. Designs can often avoid that: a phased roll-out gives everyone the programme eventually, and an encouragement design randomly offers extra help to enrol without stopping anyone else from joining.

Qualitative and theory-based evaluation

Numbers tell you how much changed; people tell you why. Interviews, group discussions, observation and case studies explain what worked, for whom and in what circumstances, and they catch effects nobody planned, good or bad.

Theory-based approaches go further. Contribution analysis, for example, takes the project’s theory of change and tests each link with whatever evidence is available, to judge whether and how the project contributed to a result. For small or new programmes, where a comparison group isn’t possible, a careful theory-based evaluation is often the most useful and honest choice.

Qualitative work needs rigour too: choose whom to interview for a reason and say what it was, record and analyse systematically, look for evidence that contradicts your conclusions, and check findings against other sources. This cross-checking is called triangulation.

Sample sizes in plain words

How many households or children you need depends on what you want to know.

Detecting a change between two groups, as an impact evaluation must, needs careful power calculations. J-PAL’s guide to them makes four points worth remembering:

  • Smaller effects need much bigger samples. Halving the effect you want to detect quadruples the sample you need.
  • Clustering matters. If you assign whole schools or villages, the number of schools or villages matters a great deal, often more than the number of children in each.
  • Take-up matters. If only a quarter of people offered the programme take it up, you need to offer it to 16 times as many people to detect the same effect as with full take-up.
  • People drop out. Plan for families who move or can’t be found at the endline.

By convention, studies aim for an 80% chance of detecting a real effect, at a 5% significance level. If your budget can’t reach the sample this implies, say so, and choose a design that answers a question you can afford.

Choosing an evaluator, budgets and ethics

Independence. For a credible evaluation, the evaluator shouldn’t be the team that ran the project or anyone paid on its results. Agree at the start that findings will be reported as they are, good or bad. The detailed terms of reference for a mandatory CSR assessment are in our guide to impact assessment under the CSR rules.

Skills. Look for a field team that speaks the local language, statistical skills for the design you need, qualitative skills, sound ethics and data protection practice, and examples of past reports, including some with disappointing findings.

Budget. The cost is driven by the sample size, the number of survey rounds, distances and the season, any tests or measurements, and the time for analysis and writing. Each extra round repeats most of the field cost, so plan the baseline and endline together from the start. Under the CSR Rules, a company can count the cost of an impact assessment as CSR spending up to 2% of the year’s total CSR expenditure or ₹50 lakh, whichever is higher.

Ethics.

  • Ask for informed consent, in the local language, and make clear that refusing won’t affect anyone’s access to services.
  • For children, ask a parent or guardian and the child, and follow your safeguarding policy.
  • Don’t deny a comparison group anything it would otherwise get; where you can, offer it the programme later.
  • Protect personal data under the Digital Personal Data Protection Act, 2023 and its 2025 Rules, and remove names before sharing data with funders.
  • For health measurements, such as children’s weight or anaemia tests, use trained staff, refer anyone who needs care, and seek ethics committee review.
  • Share the findings with the community, not just the funder.

Questions people ask

What is a baseline survey?

It is a measurement of a project’s key indicators before the work starts, such as the share of children reading at grade level or households with safe water. It is the starting point against which change is measured. Without one, a project can show what it did, but not what changed.

What is the difference between a baseline and an endline?

The baseline is measured before the project begins and the endline at its end, using the same indicators, questions and methods. Comparing the two shows how much changed. To know how much of that change the project caused, you also need a comparison group or another way of estimating what would have happened anyway.

What is a counterfactual in evaluation?

It is what would have happened to the same people or places without the project. It can’t be observed directly, so evaluations estimate it, for example with a comparison group of similar villages or schools. The project’s effect is the difference between what happened and the counterfactual.

Do we need a randomised controlled trial?

Usually not. Randomised trials are most useful when a big decision depends on the answer, the programme is stable, a lottery is fair and the sample is large enough. Many programmes are better served by a comparison group design, a phased roll-out or a careful theory-based evaluation.

How big should a baseline sample be?

To estimate a share to within 5 percentage points with 95% confidence, a simple random sample needs about 385 households. Sampling villages first means you need more, depending on how alike households in the same village are, and detecting a change between groups needs more again. Ask a statistician to calculate it for your design.

What if we didn’t do a baseline?

Rebuild what you can from records made at the time, such as school registers, health records and enrolment forms, and use official data for the area. You can also measure a comparison group now and ask carefully worded recall questions. Be honest about the limits in your report, and take a proper baseline for the next phase.

Sources

  1. Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development (second edition, 2023) · OECD
  2. Introduction to randomized evaluations · J-PAL
  3. Quick guide to power calculations · J-PAL
  4. Companies (CSR Policy) Amendment Rules, 2022: impact assessment cost limit (First Notes, 19 October 2022) · KPMG in India

Go deeper in the Academy

Is your CSR work worth a story?

The CSR Desk interviews CSR heads, and our weekly mailer reaches more than 50,000 readers.

Be seen for your work
  • Impact and reporting

    Impact assessment under the CSR rules

    Larger companies must have an independent agency assess the impact of their bigger CSR projects. Here is exactly who is covered, the cost limit, the Social Stock Exchange exemption, and how to write terms of reference that produce a useful report.

    Guide · 10 min read

  • Impact and reporting

    Choosing indicators that mean something

    An indicator is what you measure to know whether change is happening. Choose well and your reports show real results; choose badly and you end up counting kits and saplings. Here is how to choose, with examples by sector.

    Guide · 10 min read

  • Impact and reporting

    Needs assessment: start with the community

    Good programmes start from what a community needs, not from what a funder wants to build. Here is how to find out, using India’s public data, honest conversations and local government, and how to turn what you learn into priorities.

    Guide · 12 min read

  • Impact and reporting

    Social return on investment (SROI) and its limits

    SROI puts a rupee value on the changes a programme creates and compares it with what was spent. It can sharpen thinking about outcomes, but the headline ratio rests on many judgements. Here is how it works, and how to read one sceptically.

    Explainer · 10 min read