Frameworks

The Kirkpatrick model explained

Still the most-used evaluation framework in L&D, and still designed for a classroom in 1959. Both of those facts matter.

8 min read
In short

The Kirkpatrick model evaluates training across four levels: Reaction, Learning, Behaviour and Results. Around 78% of organisations run Level 1; roughly half attempt Level 2; very few reach Level 4. It remains the shared vocabulary of training evaluation, but it was designed for classroom delivery and returns scores rather than positions.

The four levels

Level Question Typical method Who runs it
1. Reaction Did they find it favourable and relevant? Post-session survey ~78%
2. Learning Did knowledge, skill or confidence change? Pre/post assessment ~50%
3. Behaviour Are they applying it on the job? Manager observation, performance data Few
4. Results Did a business outcome move? Cohort comparison, KPI tracking Very few

The drop-off down that column is the whole story of training evaluation. Each level is harder, slower and more contested than the one above it, and organisational appetite runs out somewhere around level two.

Level 1 — Reaction

What it measures: how learners felt about the experience.

How it is usually done: a five-question survey at the end, colloquially the "smile sheet".

What it is genuinely good for: detecting outright failure. If a session scores 2.1/5, something went badly wrong and you should look at it. Reaction data is also the only level that captures perceived relevance, which predicts whether people will engage with the next thing you send them.

Where it misleads: three ways, and they compound.

Peak-end effects. People rate an experience by its most intense moment and how it ended. Twenty minutes of checked-out boredom followed by five good minutes rates well.

Recall, not experience. A survey captures remembered feeling, not what happened moment to moment. The gap widens with session length.

It returns a score, not a position. This is the one that matters most for digital content. "3.8 out of 5" cannot tell you that everyone disengaged during the section on incident escalation. You cannot rewrite a number.

Level 2 — Learning

What it measures: change in knowledge, skill, attitude, confidence or commitment.

How it should be done: pre-test and post-test, so you measure the delta rather than the final score. A post-test alone conflates what was learned with what the learner already knew.

The refinement most people skip: a delayed retention check at 30 days. Immediate post-tests measure short-term recall, and the gap between immediate and delayed performance is usually sobering.

Worth adding: confidence weighting. A learner who is confidently wrong is a different risk from one who is uncertainly wrong, and a standard test cannot see the difference. In compliance training the confidently-wrong category is the one that causes incidents.

Level 3 — Behaviour

What it measures: whether people do anything differently at work.

Why it is hard: behaviour change is slow, observed by people who are not you, and confounded by everything else happening in the business. It typically shows at 30–90 days, by which point attribution is murky.

What actually works: manager observation against defined criteria, task performance data from systems people already use, error and incident rates in the trained domain, and time-to-competency for new starters.

The trap: self-reported behaviour change. Asking people whether they apply the training produces reliably optimistic answers and is barely better than Level 1.

Level 4 — Results

What it measures: movement in a business outcome — productivity, retention, incident rate, sales conversion.

Why almost nobody gets here: attribution. Your Q3 incident rate fell, and Q3 also had a policy change, a new starter cohort and a reorganisation. Isolating training's contribution requires either a genuine control group or a defensible isolation method.

The workable approach: cohort comparison. Phased rollouts give you a control group for free — the people scheduled for Q3 are your comparison in Q2. Control for what you can, and state plainly what you could not control for.

That last step is the one people skip, and it is the one that builds credibility. A CFO who sees you name your own confound trusts the rest.

The New World model

James and Wendy Kirkpatrick's update makes two substantive changes.

Plan backwards. Start at Level 4 — define the business result you are targeting — and design the programme and its evaluation from there, rather than delivering training and evaluating afterwards.

Required drivers. Explicit recognition that Level 3 behaviour change does not happen without reinforcement, coaching, accountability and reward in the workplace. Training alone rarely changes behaviour, and the model now says so.

The second point is the more useful one, because it reframes a common failure. When behaviour does not change, the problem is frequently the environment rather than the content, and a model that acknowledges this leads to better conversations than one that does not.

What it cannot do

Three specific limitations for anyone evaluating digital content.

It has no resolution below the programme. All four levels evaluate a course. None can tell you which twelve minutes of a forty-minute module to rewrite. For digital content, where engagement is continuously observable, that is a significant gap.

Level 1 was built for a room. In a classroom, the trainer reads the room in real time and the survey is genuinely the only after-the-fact instrument available. Online, you can observe attention directly — continuously, per section. Relying on recalled feeling when moment-by-moment behaviour is available is a strange methodological choice.

It says nothing about where the failure was. The model tells you whether each level succeeded. It has no mechanism for locating a failure within the content, which means it evaluates but does not diagnose.

The layer Kirkpatrick predates

Before Reaction, there is a question the model never had to ask because a classroom trainer could see the answer: were they attending at all, and if not, when did they stop? That is now measurable — through dwell, focus state, tab-switching and drop-off position — and it is the only layer that produces a specific fix rather than a verdict.

Using it well in 2026

Keep it as vocabulary. Everyone in your organisation already knows the four levels. That shared frame is worth more than its methodological accuracy, and replacing it costs more than it returns.

Do not use it as your data collection design. Use engagement instrumentation for the attention layer, pre/post with delayed retention for Level 2, one credible on-the-job indicator for Level 3, and cohort comparison for Level 4.

Report against it, measure beneath it. Present findings in Kirkpatrick's language because that is what stakeholders read. Collect the data using methods that postdate 1959.

Add the layer underneath. "Level 1 scored 3.8" plus "and 61% of learners disengaged during section 4, which we have now rewritten" is a far stronger report than either sentence alone — and only the second one describes a decision.

Frequently asked questions

Who created the Kirkpatrick model?
Donald Kirkpatrick, in a series of articles published in 1959 while he was working on his doctoral research at the University of Wisconsin. It was later formalised as the four levels and has been maintained and extended by James and Wendy Kirkpatrick.
Is the Kirkpatrick model outdated?
As a measurement method for digital content, largely yes — it predates e-learning by four decades and Level 1 in practice means a post-session survey. As a shared vocabulary for distinguishing 'people liked it' from 'people changed', it remains genuinely useful and everyone in the room already knows it.
What is the difference between Kirkpatrick and Phillips?
Phillips adds a fifth level converting Level 4 business results into monetary value, plus a structured process for isolating training's contribution from other factors. The isolation step is the substantive addition — without it, dividing benefit by cost is not an ROI calculation.
Why do so few organisations reach Level 4?
Because attribution is genuinely hard. Isolating a training programme's effect on a business outcome from everything else that changed in the same period requires either a control group or a defensible isolation method, and most organisations have neither the data nor the appetite.

See where your content loses people

Book a walkthrough and we will show you the engagement data on your own content.

Request a demo
Get Started

See what completion rates can't tell you

Find out exactly where your content works, where it fails, and what disengagement looks like before people leave.

Request a Demo