You cannot evaluate everything at Level 3 and Level 4. It sounds obvious when stated plainly, yet much of the discourse around learning impact proceeds as if every form of learning should result in observable behaviour change and measurable organisational outcomes. The expectation is applied broadly, often uncritically, and frequently without regard for context.
Consider a simple reality. Not everyone who completes a university degree knows exactly how they will apply it. Some do, certainly, but many do not. A degree expands knowledge, builds capability, and opens pathways, yet its application is often flexible, deferred, indirect, or still emerging. We accept this without hesitation in formal education. There is no insistence that every graduate demonstrate immediate behavioural change tied to a defined performance outcome. The value of the learning is recognised as foundational, enabling, and, in many cases, preparatory.
Yet in organisational learning, we routinely impose a different standard. We expect immediate transfer. We expect demonstrable behavioural change. We expect measurable results. Frameworks such as the Kirkpatrick Model are not inherently flawed in this regard. The difficulty lies in how they are interpreted, particularly at Levels 3 and 4, where evaluation moves beyond the learning environment and into the realities of work, performance, and organisational systems. What is rarely questioned is whether those levels are appropriate in the first place.
There is often the assumption that all learning should translate directly into behaviour change, and that behaviour change should, in turn, lead to measurable organisational results. It is a neat and reassuring model, but it is frequently disconnected from how organisations actually function. Some learning is indeed tightly linked to defined roles, processes, and performance expectations. In such cases, Level 3 and Level 4 evaluation is both appropriate and necessary. However, a significant proportion of organisational learning does not operate in this way.
Organisations increasingly invest in new and emerging knowledge areas not because they have fully articulated use cases, but because they recognise the need to build awareness and capability ahead of change. In these contexts, behavioural expectations are not yet stable, performance indicators are not yet agreed, and success itself is often still being defined.
Even in more established domains, evaluation frequently assumes a level of organisational clarity that simply does not exist. It is not uncommon, for example, to encounter evaluation designs that seek to measure whether individuals are applying a particular model introduced during training. On the surface, the question appears straightforward. On closer examination, it becomes far less so. Which model is being referenced, and where has it been formally defined? How is it expected to be interpreted in practice, and to what extent are managers aligned in reinforcing it? In many organisations, there are no consistent answers to these questions. Models are introduced within training environments but are not operationalised in the workplace. Expectations are implied rather than explicitly defined. Managers interpret or ignore them according to their own judgement. Under such conditions, evaluation does not measure behaviour change; it measures variation in interpretation.
This is not a failure of instrumentation. It is a failure of definition. Behaviour cannot be evaluated in any meaningful way in the absence of a shared and operationalised understanding of what that behaviour looks like in practice. Without that foundation, Level 3 evaluation has nothing stable to anchor itself to, and any conclusions drawn will be, at best, partial and, at worst, misleading.
A more fundamental challenge arises when organisations invest in new or emerging knowledge areas without a clear line of sight to application. This is increasingly common. Organisations engage with new domains because they recognise their potential relevance, not because they have already embedded them into workflows or performance structures. In such cases, there are no established behavioural models, no agreed indicators of performance, and often no consensus on what effective application would even entail. To impose Level 3 and Level 4 evaluation under these conditions is to demand evidence that cannot yet exist. The result is predictable. Weak proxies are constructed, superficial indicators are introduced, and claims are made that cannot be substantiated. Evaluation becomes performative rather than informative.
It is important, therefore, to recognise that some knowledge precedes application. Not all learning is designed to produce immediate behavioural change. Some learning exists to expand conceptual understanding, to shift mental models, and to prepare individuals and organisations for future application that has not yet been fully defined. In such cases, knowledge functions as a precursor to possibility rather than as a direct driver of immediate action. Forcing this type of learning into Level 3 and Level 4 evaluation frameworks distorts both the purpose of the learning and the integrity of the evaluation itself.
There is, however, a further risk that is less frequently acknowledged. As evaluation moves into Levels 3 and 4, it becomes entangled with performance management, accountability structures, and organisational control. In some contexts, evaluation is not used to understand learning, but to justify decisions about individuals. It becomes a mechanism through which managers are pressured, employees are assessed, and, at times, consequences are imposed. When evaluation is used in this way, it ceases to function as a tool for insight. It becomes a tool of enforcement.
Under such conditions, the integrity of the evaluation is compromised from the outset. Data is shaped by incentives. Responses are influenced by perceived risk. Behaviour is performed for measurement rather than enacted for genuine improvement. In the worst cases, evaluation can be used to justify withholding opportunities, limiting access to benefits, or reinforcing existing hierarchies under the guise of objectivity.
This is not evaluation. It is control.
For evaluation at Levels 3 and 4 to function as intended, it must be situated within a governance framework that is transparent, proportionate, and ethically grounded. Its purpose must be clearly defined as understanding and improving systems, not policing individuals. Without this, even the most technically sound evaluation design will produce distorted results.
This does not render evaluation irrelevant. It reinforces the need to use it appropriately. Where behaviours are clearly defined, embedded in workflows, and reinforced through managerial structures, Level 3 evaluation becomes both feasible and meaningful. Where those behaviours are demonstrably linked to organisational outcomes, Level 4 evaluation can be constructed, albeit with care. In such contexts, approaches such as contribution analysis and triangulation are essential. They allow evaluators to move beyond simplistic claims of causation and instead construct credible, evidence-based arguments regarding the role that learning has played within a broader system of influences.
Even then, certainty remains elusive. Organisational environments are complex, and outcomes are rarely the product of a single intervention. Contribution analysis acknowledges this by focusing on the plausibility of influence rather than definitive proof, while triangulation strengthens the evaluation by drawing on multiple sources of evidence, methods, and perspectives. Together, they provide a disciplined way of engaging with complexity without reducing it to oversimplified metrics.
There is, at times, a concern that such structured approaches render evaluation overly rigid, reducing complex human behaviour to mechanistic frameworks. In practice, the opposite is true. Without structure, evaluation becomes anecdotal and unchallengeable, reliant on perception rather than evidence. With structure, behaviour and results become visible and open to scrutiny. The purpose of structure is not to eliminate professional judgement, but to ground it. Rigid systems attempt to force certainty where none exists. Well-designed evaluation systems recognise uncertainty and manage it with discipline.
Organisations often ask whether learning has changed behaviour or improved results. These are valid questions, but they are not always the right starting point. A more fundamental question is whether the organisation has defined the behaviour it expects, and whether it understands how that behaviour connects to performance. If the answer is unclear, then Level 3 and Level 4 evaluation will not fail because of poor design. It will fail because it has nothing stable to evaluate against.
Not all learning should be forced into Level 3 and Level 4 evaluation. Where application is undefined, evaluation must reflect that reality. Otherwise, what is produced is not insight, but artefacts that create the illusion of impact without its substance. Evaluation is not about proving that learning worked. It is about understanding what learning is doing, and, just as importantly, what it is not yet positioned to do.
