On this page

What each column actually claims

Each column of a logic model makes a different kind of claim, and mixing them is the most common defect reviewers see. The W.K. Kellogg Foundation Logic Model Development Guide, still the most widely used primer in the sector, frames the model as a picture of how the program is supposed to produce change, not as a form to complete.

Inputs are what the program consumes: staff time stated as an FTE share, volunteers, partners, facilities, materials, and money. Write them as nouns with quantities. “Community support” is not an input; two donated classrooms and twelve trained volunteer tutors are.

Activities are what the program does. Write them as verbs with frequency and duration: run 60-minute tutoring sessions twice a week for 28 weeks. An activity without a schedule cannot be costed or tested.

Outputs are the countable products of activities: students enrolled, sessions delivered, attendance achieved. Outputs prove delivery, nothing more.

Short-term outcomes are the first changes in participants: knowledge, skill, behavior, or condition, each with a named measure and a time by which it should be observable.

Longer-term outcomes are the later changes the program contributes to. The honest verb here is contributes. A 28-week tutoring program contributes to grade-level reading proficiency; it does not cause it alone.

Most published templates add a row for assumptions and external context beneath the columns. Keep it. It is where the model admits what has to stay true in the world for the chain to hold.

The four arrow tests that make or break the model

A logic model fails in the arrows, not the boxes. Teams fill five columns independently, each column looks reasonable, and the chain still does not hold. Test each arrow with one question before anything else in the proposal inherits the model.

InputsActivitiesOutputsShort-term outcomesLong-term outcomes1234CapacityDeliveryEvidenceTimeline
The five columns and the four arrow tests between them. Test 1 checks capacity, test 2 checks delivery arithmetic, test 3 checks the evidence for change, and test 4 checks the timeline and attribution of the longer-term claim.

Test 1, capacity: are the inputs enough to run the activities at this scale? This is arithmetic, not optimism. If the activities need 1,120 tutor-hours and you have twelve volunteers averaging two hours a week for 28 weeks, you have 672 hours. The model fails here quietly and often, because ambitious activities read well in a narrative.

Test 2, delivery: if the activities run as planned, do the output counts follow? Two sessions a week for 28 weeks is 56 session dates. If the outputs column says 80, the columns were filled independently. This is the easiest arrow to verify and the one reviewers check first.

Test 3, evidence: why should this dose of outputs produce the short-term outcome? This is the real causal leap. The answer can be published evidence behind the curriculum, prior results from your own records, or a stated professional rationale, but it must exist somewhere the proposal can point to. If nothing supports the arrow, shrink the outcome claim until something does.

Test 4, timeline: does the short-term change lead to the longer-term change within a window you can claim? Grant periods are short and community change is slow. State what the program contributes and over what horizon, and leave system-level change to the assumptions row if your program is one actor among many.

The CDC Program Evaluation Framework treats a clear program description as a precondition for any credible evaluation, which is exactly what a tested logic model provides: columns four and five become the outcomes your evaluation plan measures, and column one becomes the personnel and cost basis of your grant budget.

A worked model for a fictional reading program

Worked example

Cedar Bend Youth Alliance, a fictional nonprofit, models its after-school reading program

Cedar Bend Youth Alliance is a fictional organization used only to demonstrate the method. Its program: after-school reading tutoring for 40 elementary students across one school year.

ColumnEntries
Inputs0.4 FTE program coordinator; 12 trained volunteer tutors; licensed tutoring curriculum; 2 classrooms donated by the district; $48,000 program budget
ActivitiesRecruit and train tutors in August; run 60-minute tutoring sessions twice weekly for 28 weeks; hold 8 monthly family reading nights
Outputs40 students enrolled; 56 session dates delivered; 75 percent average attendance; 8 family nights held
Short-term outcomesStudents improve oral reading fluency between fall and spring benchmark tests; students report higher reading confidence on a year-end survey
Longer-term outcomesA larger share of participating students reaches grade-level proficiency on the district assessment, a change the program contributes to alongside classroom instruction
AssumptionsThe district continues providing space and benchmark data; at least 10 of 12 tutors stay through spring; families can attend evening events

Now the arrows. Test 1: sessions run in two rooms with up to 20 students each, so a session needs 6 tutors at a 1:3 or 1:4 ratio; with 12 tutors alternating, each volunteers one evening a week, which the recruitment commitment supports. Test 2: 28 weeks x 2 sessions = 56 dates, matching the outputs column. Test 3: the curriculum publisher reports fluency gains at two sessions weekly, and Cedar Bend’s pilot-year records show similar direction; the 75 percent attendance output exists precisely because the evidence assumes that dose. Test 4: grade-level proficiency is claimed as a contribution measured on the district’s own assessment, not as a result the program causes alone.

Every number in this table now has a second job: the budget prices the coordinator at 0.4 FTE, the evaluation plan measures the fluency benchmark, and the narrative describes 56 sessions. One source, many documents.

A second complete model, built for a fictional job-readiness program with an alignment walkthrough, deliberately planted misalignments, and a scoring rubric, is in the nonprofit logic model example.

Outputs are not outcomes, and reviewers know it

The fastest credibility test a reviewer runs on a logic model is whether column three leaked into column four. Attendance, materials distributed, and sessions held are delivery facts. Calling them outcomes signals that the program has not thought about change.

The sorting rule: if your team can make the number go up without any participant changing, it is an output. You can raise enrollment with better recruiting. You cannot raise a fluency benchmark that way. When a funder’s form uses different labels, and many merge outcomes and impact or say results, keep the internal distinction anyway and translate at the end.

The same discipline protects you from the opposite failure, promising outcomes with no output floor beneath them. Change needs a dose. If the evidence behind your curriculum assumes 75 percent attendance, that attendance level belongs in the outputs column as a stated requirement, and your reporting should say what happens when it is missed.

Fit the model to capacity, then to the funder

Two checks remain before the model is done.

Capacity fit. A model can pass all four arrow tests on paper and still exceed what your staff, partners, and budget can hold. Read the inputs column against reality: total effort across all programs, volunteer turnover, the partner agreement that exists only as goodwill. This is also where a logic model differs from a theory of change: the theory explains why change happens in your community at any scale, while the logic model commits one funded program, at one scale, to one delivery plan. If a funder asks for both, the logic model must be the smaller, harder-edged document, and the two must tell the same causal story.

Funder translation. Keep one internal model per program and translate it to each application. Map your labels to the funder’s labels, cut columns they do not ask for, and adopt their form when they supply one; the funder’s live instructions always win. What you should not do is redesign the program per application. The model holds the program still while the packaging changes, which is the same principle behind keeping one core narrative that individual proposals adapt.

Logic model worksheet

The five columns plus the assumptions row and all four arrow tests, as a fillable worksheet with an evidence column and a confirmed / estimated / open status field for every entry.

CSV worksheet

A tested logic model cannot guarantee an award, and no template can. What it removes is one of the most common reasons reviewers stop trusting a proposal: a causal chain that contradicts its own budget and narrative. Verify any funder-specific format requirements on the live opportunity before you translate the model into an application.