The Kirkpatrick Model in 2026: What It Gets Right and What It Misses

The Kirkpatrick Model has been the global standard for training evaluation since 1959. Used correctly; starting from Level 4 and working backward, it remains the most practically useful evaluation framework available to enterprise L&D. Used as most organisations use it; Levels 1 and 2 only, with Levels 3 and 4 perpetually deferred, it produces evaluation…


1. The Four Levels: What Each Measures and What Each Cannot

The Kirkpatrick Model defines training evaluation as a four-level hierarchy, where each successive level produces more valuable evidence of training effectiveness, and requires more effort to obtain. The framework’s elegance is its clarity. Its practical problem is that the effort required increases sharply between Level 2 and Level 3, and most organisations stop at Level 2.

the Kirkpatrick Model has been the world’s most widely used framework for measuring the impact of learning — predating most current L&D platforms by decades yet remaining the reference standard (Kirkpatrick Partners 2026)

where most organisations stop — measuring learner satisfaction and knowledge gain while Levels 3 and 4 (behaviour and results) remain unmeasured and the cascade to business impact is broken (Valamis Kirkpatrick Guide 2026)

the recommended wait before measuring Level 3 behaviour change — significantly longer than the post-training measurement window most organisations actually use (Valamis Kirkpatrick Guide 2026)

requires access to operational metrics in CRM, ERP, safety, or clinical systems that L&D rarely has standing relationships with — the structural barrier most commonly cited for non-implementation (Knowlify Kirkpatrick Guide 2026)

Key Distinction

Level 1 tells you how learners felt about the training. Level 2 tells you what they knew immediately after it. Level 3 tells you whether they changed their behaviour on the job. Level 4 tells you whether the business outcomes the training was designed to improve actually moved. Only Levels 3 and 4 answer the question that justifies the training investment. Most organisations are answering only Levels 1 and 2; which is the equivalent of measuring whether a medical treatment was pleasant and whether patients understood the dosage instructions, without checking whether the condition improved.

LevelWhat It MeasuresHow CollectedBusiness Value
1 — ReactionLearner satisfaction, perceived relevance, engagementPost-training survey (smile sheet)Low — predicts completion but not behaviour change
2 — LearningKnowledge gain, skill acquisition, confidence increasePre/post assessment, quiz, simulation performanceMedium — confirms learning occurred but not application
3 — BehaviourOn-the-job application of trained skills and knowledgeManager observation, performance review, call monitoring, operational audit, at 3–6 monthsHigh — first level where behaviour change is confirmed
4 — ResultsBusiness outcomes the training was designed to improveOperational metrics in CRM, ERP, safety, quality, or financial systemsHighest — directly answers whether the investment delivered value

2. The Cascade Break — Why Most Organisations Never Reach Level 3

The cascade from Level 1 through Level 4 is the model’s design logic — each level’s evidence builds on the previous. In practice, most organisations experience what might be called the cascade break: the four levels are treated as four disconnected measurement events because the data infrastructure, the cross-functional relationships, and the time required for Levels 3 and 4 are not in place when the programme launches.

  1. Level 1 and 2 data is automatic. Level 3 and 4 data is not. LMS platforms automatically generate completion records and assessment scores. Post-training surveys are sent automatically. Level 1 and 2 data arrives without deliberate effort. Level 3 requires a 90-day follow-up process with manager observations or performance data access. Level 4 requires relationships with the teams who own the operational metrics the training was designed to move. Neither exists without deliberate pre-programme design.
  2. The baseline must be established before the programme launches. Level 4 measurement requires a pre-programme baseline of the business metric the training is designed to improve. If that baseline is not collected before the programme starts, the post-training measurement has no comparison point. This requires designing the measurement framework before designing the content — which most L&D commissions do not require.

3. The New World Kirkpatrick Model — What Changed and Why It Matters

Jim and Wendy Kirkpatrick’s New World model makes two significant updates to the original that are directly relevant to enterprise L&D practice in 2026. First, it explicitly mandates designing from Level 4 backward, starting with the specific business outcome and working back to the behaviour and learning required to achieve it. This is a design discipline, not just an evaluation discipline. Second, it introduces the concept of required drivers, the management reinforcement behaviours, process supports, and environmental conditions that are necessary for training-produced behaviour change to persist after the learning event.

“Required drivers are the reason a training programme that produced measurably positive Level 3 data at 30 days can produce poor Level 3 data at 90 days. If the manager does not reinforce the trained behaviour, the old behaviour returns. The model that does not account for this gives L&D credit or blame for outcomes that the management environment determines.”


4. What the Model Misses — Three Legitimate Criticisms

  1. It assumes a causal chain that is often disrupted. The Kirkpatrick cascade assumes that positive Level 2 data produces positive Level 3 data, which produces positive Level 4 data. In practice, each transition is mediated by factors outside training, manager behaviour, process design, tool availability, organisational culture. Training that produces excellent Level 2 results can produce poor Level 3 results if the environment does not support application. The model identifies this risk through required drivers but does not provide a framework for diagnosing or addressing the environmental barriers specifically.
  2. It does not address the forgetting curve. The original Kirkpatrick model is silent on the decay of trained behaviour without reinforcement. A programme that produces strong Level 3 data at 30 days but produces poor data at 90 days; because no reinforcement was built in, has produced transient behaviour change, not durable capability development. This is the gap where the forgetting curve intersects with the Kirkpatrick framework, and it requires a reinforcement architecture that the model’s design does not mandate.
  3. It is most commonly applied at its least valuable levels. The structural incentive of the model; that Levels 1 and 2 are easy and Levels 3 and 4 are hard, produces an evaluation culture where organisations consistently measure the two least valuable levels and defer the two most valuable ones. This is not a model failure. It is an implementation failure that the model cannot self-correct for. Addressing it requires a cultural and commissioning change — that L&D functions are expected to demonstrate Level 3 and Level 4 evidence, not just Level 1 and Level 2 completion.

In Summary

The Kirkpatrick Model is not broken. It remains the most accessible and most widely understood framework for connecting training to business outcomes. Its logic, Reaction to Learning to Behaviour to Results; is correct, useful, and sufficiently flexible to serve any training context in any sector. The problem is not the model. It is the consistent failure to implement its most valuable levels.

The practical improvement available to most enterprise L&D functions is not a different evaluation framework. It is the discipline to establish Level 4 baselines before programmes launch, build the cross-functional relationships to access Level 3 and Level 4 data, and treat the measurement framework as a design deliverable rather than a retrospective reporting exercise. That improvement is achievable within the Kirkpatrick framework; and it would change the quality of the evidence most L&D functions can provide for their investment significantly.


Frequently Asked Questions

Q1

What are the four levels of the Kirkpatrick Model?

Level 1 Reaction measures learner satisfaction and perceived relevance. Level 2 Learning measures knowledge and skill gained. Level 3 Behaviour measures on-the-job application at 3–6 months post-training. Level 4 Results measures impact on the business outcomes training was designed to improve.


Q2

Why do most organisations stop at Levels 1 and 2?

Levels 1 and 2 data is automatically generated by LMS platforms and post-training surveys. Levels 3 and 4 require manager observation processes, access to operational metrics systems, and cross-functional relationships that must be established deliberately before the programme launches; not by default.


Q3

What does the New World Kirkpatrick Model add?

It mandates designing from Level 4 backward; starting with the business outcome and working back through behaviour and learning requirements. And it introduces required drivers: the management reinforcement behaviours and environmental conditions necessary for training-produced behaviour change to persist. Without required drivers, trained behaviour typically reverts.


Q4

What are the main criticisms of the Kirkpatrick Model in 2026?

The assumed causal chain between levels is often disrupted by environmental factors outside training. The model does not address the forgetting curve or reinforcement needed to sustain behaviour change. And it is most commonly applied at its least valuable levels — 1 and 2 — while the most valuable — 3 and 4 — are consistently deferred.


Qquench Specialists

25+ years designing enterprise training with Level 3 and Level 4 evaluation built in from the brief; because Levels 1 and 2 cannot tell you whether training changed anything in the business. We write from practice, not position papers.