The Kirkpatrick model evaluates training across four levels: Reaction, Learning, Behaviour and Results. Around 78% of organisations run Level 1; roughly half attempt Level 2; very few reach Level 4. It remains the shared vocabulary of training evaluation, but it was designed for classroom delivery and returns scores rather than positions.
The four levels
| Level | Question | Typical method | Who runs it |
|---|---|---|---|
| 1. Reaction | Did they find it favourable and relevant? | Post-session survey | ~78% |
| 2. Learning | Did knowledge, skill or confidence change? | Pre/post assessment | ~50% |
| 3. Behaviour | Are they applying it on the job? | Manager observation, performance data | Few |
| 4. Results | Did a business outcome move? | Cohort comparison, KPI tracking | Very few |
The drop-off down that column is the whole story of training evaluation. Each level is harder, slower and more contested than the one above it, and organisational appetite runs out somewhere around level two.
Level 1 — Reaction
What it measures: how learners felt about the experience.
How it is usually done: a five-question survey at the end, colloquially the "smile sheet".
What it is genuinely good for: detecting outright failure. If a session scores 2.1/5, something went badly wrong and you should look at it. Reaction data is also the only level that captures perceived relevance, which predicts whether people will engage with the next thing you send them.
Where it misleads: three ways, and they compound.
Peak-end effects. People rate an experience by its most intense moment and how it ended. Twenty minutes of checked-out boredom followed by five good minutes rates well.
Recall, not experience. A survey captures remembered feeling, not what happened moment to moment. The gap widens with session length.
It returns a score, not a position. This is the one that matters most for digital content. "3.8 out of 5" cannot tell you that everyone disengaged during the section on incident escalation. You cannot rewrite a number.
Level 2 — Learning
What it measures: change in knowledge, skill, attitude, confidence or commitment.
How it should be done: pre-test and post-test, so you measure the delta rather than the final score. A post-test alone conflates what was learned with what the learner already knew.
The refinement most people skip: a delayed retention check at 30 days. Immediate post-tests measure short-term recall, and the gap between immediate and delayed performance is usually sobering.
Worth adding: confidence weighting. A learner who is confidently wrong is a different risk from one who is uncertainly wrong, and a standard test cannot see the difference. In compliance training the confidently-wrong category is the one that causes incidents.
Level 3 — Behaviour
What it measures: whether people do anything differently at work.
Why it is hard: behaviour change is slow, observed by people who are not you, and confounded by everything else happening in the business. It typically shows at 30–90 days, by which point attribution is murky.
What actually works: manager observation against defined criteria, task performance data from systems people already use, error and incident rates in the trained domain, and time-to-competency for new starters.
The trap: self-reported behaviour change. Asking people whether they apply the training produces reliably optimistic answers and is barely better than Level 1.
Level 4 — Results
What it measures: movement in a business outcome — productivity, retention, incident rate, sales conversion.
Why almost nobody gets here: attribution. Your Q3 incident rate fell, and Q3 also had a policy change, a new starter cohort and a reorganisation. Isolating training's contribution requires either a genuine control group or a defensible isolation method.
The workable approach: cohort comparison. Phased rollouts give you a control group for free — the people scheduled for Q3 are your comparison in Q2. Control for what you can, and state plainly what you could not control for.
That last step is the one people skip, and it is the one that builds credibility. A CFO who sees you name your own confound trusts the rest.
The New World model
James and Wendy Kirkpatrick's update makes two substantive changes.
Plan backwards. Start at Level 4 — define the business result you are targeting — and design the programme and its evaluation from there, rather than delivering training and evaluating afterwards.
Required drivers. Explicit recognition that Level 3 behaviour change does not happen without reinforcement, coaching, accountability and reward in the workplace. Training alone rarely changes behaviour, and the model now says so.
The second point is the more useful one, because it reframes a common failure. When behaviour does not change, the problem is frequently the environment rather than the content, and a model that acknowledges this leads to better conversations than one that does not.
What it cannot do
Three specific limitations for anyone evaluating digital content.
It has no resolution below the programme. All four levels evaluate a course. None can tell you which twelve minutes of a forty-minute module to rewrite. For digital content, where engagement is continuously observable, that is a significant gap.
Level 1 was built for a room. In a classroom, the trainer reads the room in real time and the survey is genuinely the only after-the-fact instrument available. Online, you can observe attention directly — continuously, per section. Relying on recalled feeling when moment-by-moment behaviour is available is a strange methodological choice.
It says nothing about where the failure was. The model tells you whether each level succeeded. It has no mechanism for locating a failure within the content, which means it evaluates but does not diagnose.
Before Reaction, there is a question the model never had to ask because a classroom trainer could see the answer: were they attending at all, and if not, when did they stop? That is now measurable — through dwell, focus state, tab-switching and drop-off position — and it is the only layer that produces a specific fix rather than a verdict.
Using it well in 2026
Keep it as vocabulary. Everyone in your organisation already knows the four levels. That shared frame is worth more than its methodological accuracy, and replacing it costs more than it returns.
Do not use it as your data collection design. Use engagement instrumentation for the attention layer, pre/post with delayed retention for Level 2, one credible on-the-job indicator for Level 3, and cohort comparison for Level 4.
Report against it, measure beneath it. Present findings in Kirkpatrick's language because that is what stakeholders read. Collect the data using methods that postdate 1959.
Add the layer underneath. "Level 1 scored 3.8" plus "and 61% of learners disengaged during section 4, which we have now rewritten" is a far stronger report than either sentence alone — and only the second one describes a decision.
Frequently asked questions
Who created the Kirkpatrick model?
Is the Kirkpatrick model outdated?
What is the difference between Kirkpatrick and Phillips?
Why do so few organisations reach Level 4?
See where your content loses people
Book a walkthrough and we will show you the engagement data on your own content.