Kirkpatrick vs Phillips vs Brinkerhoff – Which Training Evaluation Framework Is Right for Enterprise L&D

Kirkpatrick, Phillips ROI, Brinkerhoff’s Success Case Method, and CIRO are not competing alternatives. They are complementary frameworks with distinct purposes, and the choice between them is not ideological. It is practical. Most mature enterprise L&D functions run Kirkpatrick as the evaluation spine and layer one of the others for a specific measurement gap. Here is…


1. The Four Frameworks: What Each Actually Does

The reason most enterprise L&D teams struggle to choose between evaluation frameworks is that they are usually described as if they answer the same question. They do not. Each framework answers a different question about training and the right framework for any programme is determined by which question its stakeholders are actually asking.

FrameworkCore Question It AnswersWhat It ProducesWho Needs It
Kirkpatrick (4 Levels)Did the training work at four increasing levels of impact?Reaction data, learning evidence, behaviour change assessment, business results. The full evaluation chain from learner experience to business outcome.Every L&D function evaluating any programme. Kirkpatrick is the default evaluation structure for enterprise training globally.
Phillips ROI (5 Levels)Was the training worth the investment in financial terms?Kirkpatrick Level 4 results converted to monetary value, compared against total programme cost, producing a financial ROI percentage. CFO-grade business justification.High-investment programmes — enterprise leadership development, large compliance platforms, major capability transformation — where CFO-facing financial justification is required.
Brinkerhoff Success Case MethodWhat made the training succeed for some participants and not others?Qualitative interview data from the top 10% and bottom 10% of performers, identifying what organisational conditions, management behaviours, and individual characteristics drove the difference.Programmes where quantitative scores are incomplete evidence — leadership development, behavioural skills training, change management — and where understanding the why of success and failure is as important as measuring it.
CIRO (Context, Input, Reaction, Outcome)Is the programme design right before it runs?Front-loaded evaluation of the problem context and programme design quality before measurement of outcomes. Prevents the failure mode of evaluating a poorly designed programme and attributing weak results to learners rather than design.High-investment programmes at design stage where getting the brief right before launch is more valuable than measuring a poor design after the fact.

Key Distinction

These frameworks are not ranked in sophistication. Kirkpatrick is not simpler than Phillips ROI — it is different. CIRO is not more advanced than Brinkerhoff; it operates at a different stage of the programme lifecycle. The choice is not which framework is best. It is which question your stakeholders are asking, and which framework answers it most directly.


2. Kirkpatrick — The Universal Spine

Kirkpatrick is the default evaluation structure for enterprise L&D globally because it provides the most complete picture of training impact across the full chain from learner experience to business outcome.

The four levels — Reaction, Learning, Behaviour, Results — are not optional stages to be selected based on resource availability. They are a causal chain. Learner reaction affects learning. Learning affects behaviour change. Behaviour change produces results. Measuring only Level 4 results without measuring Levels 2 and 3 means not knowing why results moved or did not, which means the next iteration of the programme is as likely to repeat the failure as to address it.

L1

Reaction — Did learners find the training relevant and valuable?

Measured by post-training surveys. Often called “smile sheets.” Useful as a leading indicator of engagement and perceived relevance. Not sufficient as a standalone measure of programme effectiveness. High satisfaction scores on poorly designed training are common — they measure the production quality of the experience, not its capability change impact.

L2

Learning — Did learners acquire the knowledge, skill, or attitude the training was designed to produce?

Measured by pre/post assessments, skills demonstrations, or scenario-based assessment. This is the level where most compliance training measures — and stops. Knowledge acquisition is necessary but not sufficient. It does not prove the knowledge transferred to behaviour.

L3

Behaviour — Did learners apply the trained behaviour in their work context?

Measured by manager observation, operational audit, call quality monitoring, or performance management data — four to eight weeks after training. This is the level most organisations do not reach. It requires data from operational systems that L&D rarely accesses and measurement timing that extends beyond the immediate post-training window.

L4

Results — Did the training produce the business outcome it was designed to improve?

Measured by the operational metric the programme was designed to move — incident rate, win rate, defect rate, compliance breach rate. This level requires baseline data collected before training and cross-functional data access to the operational systems where results are recorded. It cannot be retrospectively assembled after the programme concludes.

The practical rule for Kirkpatrick implementation: start at Level 4 and design backward. Define the business result the programme must produce first. Identify the behaviours at Level 3 that produce that result. Design the learning at Level 2 that produces those behaviours. Ensure the experience at Level 1 engages learners in the Level 2 design. Most L&D functions design forward — which produces programmes that satisfy Level 1 without guaranteeing any connection to Level 4.


3. Phillips ROI — When the CFO Is in the Room

The Phillips ROI Model extends Kirkpatrick with a fifth level: the financial return on the training investment. It converts Level 4 results into monetary value, subtracts programme costs, and produces an ROI percentage that speaks the language of finance rather than learning.

“The Phillips model doesn’t only collect data to find if the training worked or not — it evaluates the WHY behind the success or failure, adding qualitative feedback to the financial data process to help organisations improve their training programmes.” The financial calculation is the output. The qualitative understanding of what produced it is what makes the calculation useful for programme improvement.

The formula is: ROI% = (Net Programme Benefits ÷ Programme Costs) × 100. If a sales capability programme costs $50,000 and produces a measurable $200,000 increase in revenue directly attributable to improved advisory skills, the ROI is 300%.

The challenge is isolation, establishing that the outcome improvement was produced by training rather than other variables (market conditions, product changes, management turnover). Phillips ROI requires comparison groups, trend analysis, or expert estimation to isolate the training contribution. This is the most demanding aspect of the framework — and the most important, because without it, the financial calculation is correlation rather than causation.

When to use Phillips ROI: high-cost leadership development where the board is scrutinising learning investment; enterprise compliance platforms where the cost of non-compliance is quantifiable; sales capability programmes where revenue attribution is possible; and any programme where L&D is defending its budget against a CFO who wants financial evidence.


4. Brinkerhoff and CIRO — When Depth and Design Quality Matter More Than Financial Return

Brinkerhoff’s Success Case Method and CIRO fill two specific gaps that Kirkpatrick and Phillips do not address — and they are the right frameworks for specific programme types and evaluation questions that the more widely used models cannot answer adequately.

Brinkerhoff — Why did training succeed for some and not others?

Brinkerhoff identifies the top 10% and bottom 10% of programme participants through quantitative screening, then conducts deep qualitative interviews with both groups. The goal is not to measure average outcomes — it is to understand what conditions, behaviours, and support structures enabled the top performers and what barriers prevented the bottom performers from applying training.

This framework is particularly valuable for leadership and behavioural skills programmes where the success factors are contextual — where manager support, organisational culture, and individual circumstance determine whether training transfers, and where aggregate scores conceal the variation that matters most to programme improvement.

CIRO — Is the programme design right before it runs?

CIRO front-loads evaluation to the programme design stage. By evaluating the operational Context (what problem is being addressed), the programme Input (whether the design is adequate to address it), early Reaction data, and eventual Outcome — it builds quality assurance into the design process rather than measuring outcomes after a potentially flawed design has been delivered.

The failure mode CIRO prevents: evaluating a poorly designed programme and attributing weak results to learner motivation or management follow-through rather than the programme design that produced them. For high-investment programmes where redesign after delivery is expensive, CIRO’s front-loaded quality assurance is the more economical evaluation investment.


In Summary

Kirkpatrick is the evaluation spine. Run it on every programme of consequence, starting at Level 4 and designing backward to Level 1. Add Phillips ROI when the funding conversation requires financial justification that Kirkpatrick Level 4 alone does not provide. Add Brinkerhoff when understanding why training succeeded for some and failed for others is as important as the aggregate outcome score. Use CIRO when programme design quality must be assured before delivery rather than diagnosed after it.

The most mature enterprise L&D evaluation approaches are not those that have chosen one framework. They are those that use the right framework for the right programme, treating evaluation model selection as a deliberate programme design decision, not as a post-delivery reporting format.


Frequently Asked Questions

Q1

What is the difference between Kirkpatrick and Phillips ROI?

Kirkpatrick evaluates training across four levels — Reaction, Learning, Behaviour, and Results — proving whether training worked. Phillips ROI adds a fifth level that converts Level 4 results into financial terms using a cost-benefit analysis, producing an ROI percentage for CFO-facing financial justification. Kirkpatrick proves whether training worked. Phillips ROI proves whether it was worth the investment in financial terms a CFO will accept.


Q2

When should L&D teams use Brinkerhoff’s Success Case Method?

Brinkerhoff is right when quantitative metrics alone do not capture what made training succeed or fail. By studying the top 10% and bottom 10% of participants through qualitative interviews, it identifies what organisational conditions, management behaviours, and individual characteristics drove the difference. Particularly valuable for leadership and behavioural skills programmes where success factors are contextual and aggregate scores conceal the variation that matters most for improvement.


Q3

What is the CIRO evaluation model and when is it most useful?

CIRO (Context, Input, Reaction, Outcome) front-loads evaluation before the programme runs — assessing the problem context and design quality before measuring outcomes. It prevents the failure mode of evaluating a poorly designed programme and attributing weak results to learners rather than design. Most useful in the planning stage of high-investment programmes where getting design right before launch is more valuable than measuring a poor design after it.


Qquench Specialists

25+ years selecting and designing evaluation frameworks for enterprise training programmes across Fortune 100 clients globally. We write from practice, not position papers.