Why “the feedback was great” is not enough anymore

Main takeaways

  • Credible evidence starts with what participants did during the session. Not what they reported feeling about it afterwards.
  • Simulation-generated performance data is structurally different from a post-session survey. It captures what people did under pressure, not how they felt about it.

What standard training metrics actually measure

Post-session surveys were designed to capture participant reaction.

Whether the facilitator was good.
Whether the content felt relevant.
Whether participants would recommend the programme.

Donald Kirkpatrick formalised this in 1959 as Level 1 of his four-level evaluation model. His model goes further than just reactions:

  1. Reaction - did participants respond positively?
  2. Learning - did they acquire the knowledge?
  3. Behaviour - did they apply it on the job?
  4. Results - did it produce a business outcome?

Most training measurement stops at Level 1.

Some reaches Level 2.

Very little reaches Level 3 or 4.

The reason is practical. Level 1 is easy to measure. A form goes out while the room is still warm. Level 3 requires structured observation after the fact. Line managers need to agree upfront on what behavioural change looks like. Follow-up runs at 30 and 60 days.

Most programmes do not build that infrastructure. The survey is what gets returned, and the survey is what gets reported.

The problem is not that satisfaction data is useless. The problem is that it is being used to answer a question it was not built to answer.

A business unit head asking whether the training investment is producing returns is asking a Level 4 question. Responding with Level 1 data does not address it. It signals that the right measurement infrastructure was never built.

What a business unit head is actually evaluating

A business unit head does not think about training in terms of learning outcomes.

They think about risk and return.

On the return side

The questions are practical:

  • Did the investment change how people work?
  • Did that change produce a measurable result?
  • Did people reach independent performance faster?
  • Did errors go down?
  • Did new hires become productive in the expected time?

These are the metrics that connect to revenue, cost, or risk. Survey scores do not.

On the risk side

In regulated industries, the question becomes sharper.

A compliance programme that produced good feedback scores but did not change behaviour is not a neutral outcome. It is a liability.

A regulator asking what the firm did to ensure staff applied the relevant framework will not accept a completion certificate as proof of behaviour change.

A satisfaction score is not a defence either.

What simulation-generated data makes possible

A classroom session or e-learning module usually produces attendance records. If a post-session assessment runs, it may also produce a score. Neither tells you how the participant behaved under job conditions.

A Finsimco simulation produces data during the session itself. Participants make sequential decisions under time pressure, with other participants in counterpart roles. Every decision is recorded.

That makes it possible to see:

  • where participants hesitated;
  • where they moved quickly;
  • where they made the wrong call;
  • where they corrected it;
  • where they failed to adjust;
  • and how they performed when the situation changed.

That is observable behaviour under conditions that approximate the real job. Not a self-report of how confident they felt afterwards.

A frame for the internal conversation

The internal champion does not need to explain learning science to a business unit head. They need to reframe the question. The standard frame is:

“We ran a training programme and participants rated it highly.”

That answers a question the business unit head was not asking.

A more useful frame has three parts.

1. Name the measurement gap

The current metrics capture participant reaction. They do not capture behaviour change or business impact. That is a well-documented gap in L&D evaluation, not a criticism of the programme itself.

2. Name the risk

In financial services, training that cannot demonstrate behaviour change is a liability in its own right. Re-training carries a cost. Regulatory remediation carries a cost. Errors made by people who completed the programme but could not apply it when it mattered also carry a cost.

3. Name what better evidence looks like

Performance data tied to specific decisions and skills is the starting point for behavioural evaluation.

A pilot with a defined group and agreed success metrics is a lower-risk way to test that than a full procurement commitment.