On this page

Proportionate is the standard, not the fallback

Small organizations often write evaluation sections apologetically, as if the real standard were a control group and their plan were a compromise. The opposite is true in the published guidance. The CDC Program Evaluation Framework makes utility and feasibility explicit standards: an evaluation is judged by whether its results get used and whether it can actually be carried out, not by methodological ceremony. The W.K. Kellogg Foundation Evaluation Handbook goes further and frames evaluation primarily as a learning tool for the organization running the program.

That reframing changes what you write. A proportionate plan that names who collects what, when, and what decision the data informs is a complete answer to the evaluation question on a foundation application. A borrowed university-style design with no one to run it is not a stronger answer; it is a visible risk.

The plan also has to respect a boundary: more measurement is not automatically better, because every instrument costs staff time and asks something of participants. Collect only what you will use, and only what you can protect.

Write questions before you pick metrics

The template starts with evaluation questions, not indicators, because a metric without a question is unfalsifiable filler. A good evaluation question is one a board member would actually ask, or one whose answer would change how you run the program next year. Two or three are enough for a small program.

Questions come straight out of the program’s causal chain. If you built a logic model, columns three through five already contain them: did we deliver the dose we planned, did the short-term change appear, and is there early evidence of the longer-term contribution? A plan grounded this way cannot drift into measuring things the program never promised.

Then, and only then, pick one indicator per question. The discipline of one indicator each is deliberate: it forces the team to choose the measure it trusts most, and it keeps the collection burden inside what a part-time coordinator can sustain.

Four tests every indicator must pass

Run every candidate indicator through this table before it enters the plan. An indicator that fails any test either gets repaired or gets cut.

TestThe questionFails when
DefinedCould two staff members count this the same way without discussing it?“Improved wellbeing” with no instrument, or “served” with no definition of served
FeasibleCan current staff collect it with current tools, consent, and agreements?The indicator needs school district data and no data-sharing agreement exists
RelevantDoes it answer one of the written evaluation questions?It is tracked because a past funder asked for it once
InterpretableWill you know what a good result looks like when you see the number?No baseline, no target, and no comparison point of any kind

The interpretable test deserves one caution: do not invent precision to pass it. If the program has never measured this indicator, the honest plan sets year one as the baseline year and states that a target will be set from it. A target conjured without a starting condition is the kind of number that turns into an accountability problem in the grant report.

The data collection reality check

Before the plan is final, walk each indicator through one concrete week of program operation and answer four questions in writing.

Who collects it, by name or role? “The team” collects nothing. If the answer is the same coordinator who runs sessions, the tool must fit inside session routines, like a sign-in sheet that feeds a spreadsheet.

What tool captures it? Name the actual artifact: the benchmark score report, the five-question paper survey, the attendance spreadsheet. If the tool does not exist yet, creating it is a task in the program timeline.

When, and how often? Match timing to when change can plausibly be observed. Fluency gains do not appear in October; attendance problems do. Frequent operational indicators, rare outcome indicators.

What protects participants? State where data lives, who can see it, and what consent covers. Never collect identifiable data you have no plan to protect, and never promise anonymity a spreadsheet full of names cannot deliver.

Two method cautions belong in this check. First, qualitative material: interviews and open responses are legitimate evidence for a defined learning question, but a handful of memorable comments is not representative impact, and the plan should label quotes as illustrative. Second, the evaluation’s cost is real work: hours for collection, entry, and analysis belong in the grant budget as a small work breakdown, not as a generic percentage pasted onto the total.

A worked plan for a fictional tutoring program

Worked example

Cedar Bend Youth Alliance, a fictional nonprofit, plans evaluation for its reading program

Cedar Bend Youth Alliance is a fictional organization used only to demonstrate proportionate scale. The program: after-school reading tutoring for 40 elementary students, twice weekly across 28 weeks, run by a 0.4 FTE coordinator and volunteer tutors.

QuestionIndicatorSource and toolTimingOwnerUse
Did students attend enough to expect change?Percent of enrolled students attending at least 75 percent of sessionsSession sign-in sheets entered into a spreadsheetMonthlyProgram coordinatorBelow-threshold months trigger family outreach within two weeks
Did reading fluency improve?Change in words-correct-per-minute between fall and spring district benchmarksDistrict benchmark reports, shared under an existing data agreementSeptember and MayProgram coordinatorReported to the funder; informs next year’s session dose
What did families observe at home?Tallied responses and themes from a five-question paper surveySurvey distributed at the December and May family nightsTwice yearlyExecutive directorQuotes labeled illustrative in reports; themes inform family night content

Three features make this plan credible rather than impressive. Every indicator passes all four tests, including feasibility: the only external data source has an agreement already in place. The collection burden is roughly two hours a month plus two benchmark pulls, which a 0.4 FTE coordinator can absorb. And each row ends in a use, so no data is collected into a void. The stated limitation, written directly into the plan: no comparison group, so fluency gains are reported as change among participants, not as proof the program caused the change.

Evaluation plan template

Seven sections: program summary, evaluation questions, the four-test indicator table, the data collection plan with owners and privacy handling, analysis, use of results, and limitations.

Markdown template

Say how you will use the results

The weakest sentence in most evaluation sections is a promise of continuous improvement with no mechanism behind it. Replace it with named decisions: which meeting reviews the data, what threshold triggers what action, and what goes to the funder when. The Kellogg handbook’s core argument is that evaluation exists to inform exactly these decisions, and a funder reading a named decision process can tell the plan is real.

Include the uncomfortable branch too: what happens if results are weaker than expected. The honest answer, that the team will examine dose and delivery first and adjust the model, reads far better than silence, because every experienced reviewer knows weak first-year results are common.

A proportionate plan cannot guarantee an award or a particular result; it guarantees that whatever happens, you will know, and can say so with evidence. Check the funder’s live instructions for required measures or reporting formats before finalizing, and give the pre-submission review one specific job: confirm the evaluation plan measures exactly the outcomes the narrative promises.