Assessment proxy
Simulation Can Show Transfer, but It Cannot Inherit the Job
A well-designed simulation can reveal decisions that a quiz never reaches. It still needs an explicit argument about fidelity, scoring and what the real workplace adds.
What the wrapper says
Assessment makes a decision from observed evidence, while simulation training creates controlled scenarios whose fidelity determines what that evidence can claim about work.
The all-clear
Simulation is useful for safe practice, rare events and observable decision-making that a quiz or incident review cannot safely expose.
What is durable
The ability to recognise conditions, choose an action and verify or escalate in the real operational setting.
Thirty minutes in a simulation centre have become the phrase field ready.
A tab for every learner, in sim_assessment_results_FINAL_final2.xlsx, a pass column, and a dashboard you can show a board.
Efficient.
Also a small act of inheritance law that no simulation ever agreed to.
The scenario was a controlled representation.
The job is the job.
A simulation can show a decision under stated conditions. It cannot inherit the whole workplace.
That is not a complaint about simulation.
It is why simulation is useful.
The pass starts answering the wrong question
There is a real performance problem underneath.
A ward team calls for critical-care help late.
The on-call team arrives without a shared account of what has happened.
An educator and a clinical expert spend half an hour turning “improve communication” into two objectives.
Within five minutes of a defined deterioration, the bedside nurse states concern and requests senior review.
Before transfer, the team gives a structured handover: current status, relevant history, assessment, requested action.
Good.
That is a decision problem, not a slide-deck problem.
The simulation can make those decisions visible.
At minute three, the patient’s oxygen saturation and blood pressure change. At minute seven, a relative asks whether the patient is dying. At minute nine, a junior doctor receives another request.
At minute twelve, transfer and handover happen.
Each inject has a job.
Reveal escalation.
Prioritisation.
Delegation.
Handover.
The controller records facts.
Escalation was called at six minutes.
Blood pressure was absent from the call.
A task owner was named.
The family question went unanswered for ninety seconds.
Valuable evidence.
It is not evidence that the team will make every good decision on a crowded ward.
With the actual phone system.
A real family.
A senior colleague with a previous history.
And a patient who has declined to follow the timing sheet.
The evidence is bounded because the world was bounded.
That should feel unsurprising.
Instead, a pass marker is routinely asked to become a job description with a lanyard.
Start with the decision, not the room booking
Assessment starts with a claim, and the claim has to name a decision. What will this scenario result let somebody decide, and under which conditions?
“Completed emergency scenario” is an activity record, and you already have too many of those.
“Can recognise defined deterioration, request senior review, and deliver a structured handover within the specified scenario” is a claim somebody can inspect.
The difference is a great inconvenience to a dashboard, which prefers a single green dot.
But the claim tells you how much representation your programme actually needs.
If the target is recognising a protocol, a table-top or a controlled case may be enough.
If the target is coordinating an escalation through real roles, real information and real equipment, the conditions have to expose those things.
The World Health Organization distinguishes discussion-based tabletop exercises from operations-based drills, functional exercises and full-scale field exercises.
Those are not escalating badge levels.
They are different ways of making a decision visible while managing risk.
Your simulation needs the equivalent of an assessment blueprint.
A planned starting state.
Information flow.
Triggers after correct action.
Responses to omitted action.
A stopping point.
Otherwise it is a dramatic event with a smoke machine and no argument.
Fidelity is a condition, not a compliment
People say a simulator is high fidelity as though it has been promoted.
Fidelity is more useful treated as a statement about what the representation preserves, and what it does not.
Equipment orientation is not a polite preamble.
If a nurse cannot find the emergency call number, or a doctor cannot locate the observation chart, or a team cannot find the oxygen source, then later failure is ambiguous.
It may be judgement.
It may be that the room does not work like the work.
That is psychological fidelity doing operational work.
The fiction contract makes the boundary explicit.
Participants engage with the represented world.
Facilitators disclose material limitations, and do not turn those limitations into trick questions.
Ten minutes of prebrief for six participants can establish roles, equipment, the time limit, the stop phrase, recording boundaries, and what the manikin cannot do.
That is not ceremony.
It is part of whether an observed action can mean what the rubric later says it means.
Aviation makes the distinction unusually hard to ignore.
The FAA’s simulator framework separates flight training devices from full flight simulators across defined levels, with qualification attached to specific aircraft models, configurations, performance tests and permitted tasks.
A device level is not a generic ranking.
An unqualified virtual experience may be excellent practice.
It does not acquire regulatory credit because the headset looks like it has opinions about turbulence.
Every other sector merely has less obvious paperwork.
The scenario is doing real work, which is the difficulty
None of this argues for returning to a quiz with twenty questions and one suspiciously long answer option.
Simulation can safely expose people to rare, consequential events.
It can make an escalation decision, an information handoff, a prioritisation choice and a team’s division of work observable, in ways selected response cannot reach.
It can also reveal a system defect.
When the emergency call number is missing from a new ward phone display, the useful finding is a facilities request. Not a remedial course assigned to everyone who failed to discover it telepathically.
The healthcare simulation standards make the traceability principle explicit. A defined practice need connects to measurable objectives, to scenario events, to an observation method, to feedback, and to a decision.
The scenario, the rubric and the debrief are one production unit. Even when three teams own them and the folder is called Simulation_Assets_Archive_DO_NOT_MOVE.
Debriefing is what converts an observed episode into practice.
Twenty minutes of debrief first reconstructs the facts, then examines two or three objectives, then rehearses the replacement.
Name the escalation phrase.
Designate one caller.
Read the last two observation sets during handover.
Then, within 48 hours, the educator sends a de-identified summary.
Two strengths.
Two barriers.
Owners.
A due date.
That is how a staged event leads to a changed work condition.
The simulation is not pretending to be the job.
It is helping the job become safer and more observable.
A consequential score needs a chain of custody
The trouble starts when the result decides progression, credentialing or employment.
Then four controls matter.
Standardisation.
Rater preparation.
Scoring rules.
Appeal.
Standardisation fixes your starting state, your information, your role-player responses, your time window and your allowable prompts.
It does not promise an identical emotional experience.
It stops material differences being left to facilitator whim.
Raters apply the same rubric to sample performances and reconcile disagreement before any high-stakes use.
Every level has to name something a rater can see. “Good judgement” is an aspiration, not a scoring dimension.
Critical errors may be conjunctive: every specified safety criterion has to be met. That decision belongs in the scoring rule before candidates begin, not in a discussion afterwards.
So your assessment record needs the scenario version, the simulator configuration, the instructions, the allowed materials, the time window, the accommodation, the role-player script, the rubric, the scorer materials, the incident log and the result.
Rubric_2.1 explains a later review.
“Latest final revised” explains a family of avoidable meetings.
The same boundary protects the learners and the programme.
A candidate who misses an unfamiliar control may have a real deficiency.
If eight of ten experienced people miss it, the scenario or its promised orientation needs investigating.
A checklist does not settle that question.
Neither does a total score, however cleanly it fits into your HRIS.
Assessment does not convert a score into a permanent adjective.
It produces a qualified claim, with evidence and limitations attached.
The simulator log has met the talent platform
This is the current problem.
Simulation produces unusually attractive machine data.
Timestamps.
Clicks.
Decisions.
Prompts.
Video markers.
Checklist rows.
Rubric dimensions.
A reassuringly decimal score.
A model can ingest all of it beside learning history, job profiles and succession data, then infer that somebody who passed Scenario 3.2 has a generally portable capability.
The model is not being reckless.
Your data contract is being coy.
Scenario data needs the same provenance as any other workforce evidence.
With extra care, because it looks so complete.
Keep the decision claim.
The scenario version.
The equipment and configuration.
The role assignment.
The inject sequence.
The administration conditions.
The rubric version.
The rater record.
The incident record.
The score.
The limitation.
Let an AI system propose a learning action, or flag a question.
Do not let it silently promote a scenario-bound observation into a universal job fact.
Your analytics layer thinks timestamps are facts.
They are facts about a controlled event.
Validate transfer where the work happens
The durable object is not a simulation pass.
It is the ability to recognise the condition, choose the action, verify it, escalate it and adapt, in the real operational setting.
Which is harder to store.
Naturally your system would prefer the pass.
So your transfer claim needs a second evidence path.
After the session, record participation, objectives, observations, corrective actions and completion.
Test the operational change with a three-minute micro-drill where that makes sense.
Repeat a scenario after the point of prior success, and change the conditions, so the group has to apply a principle rather than recite a script.
Then observe the target behaviour in the field, and inspect whether your work environment now supports it.
The published evidence base is honest about its own limits. A systematic review of crisis-resource-management simulation found nine eligible studies.
Four measured workplace transfer.
Five measured patient outcomes.
The finding was suggestive rather than conclusive.
That is not disappointing.
That is the honest shape of the claim.
A simulation can show that a team recognised a cue, selected an action and coordinated a response inside a designed world. It can surface an instructional need, or a system barrier that deserves attention on Monday morning.
It cannot absorb the equipment, the pressure, the consequences, the local workarounds, the team history and the genuine uncertainty of the job, because the result field says PASS.
Keep the pass.
Keep its boundaries with it.
What the record establishes
- Write a decision claim that names the work, conditions, decision and consequence before selecting a simulation format.
- Select simulation fidelity and exercise form for the operational risk and the decision that the evidence must support.
- Design objectives, scenario events, observation methods, scoring rules and debrief prompts as one traceable chain.
- Retain the simulator configuration, scenario version, administration conditions, rubric and incident record with every consequential result.
- Use field observation and changed work conditions to validate transfer instead of treating a simulation pass as unrestricted readiness.
Asked in the review
- What does a simulation pass actually show?
- A simulation pass shows that a participant met the stated scoring rule in the specified scenario, with its available information, equipment, prompts, time window and rater process. It can support a claim about observable decisions under those conditions. It does not by itself establish unrestricted readiness in a real workplace, where equipment, consequences, workload, team relationships and competing priorities can differ.
- How is a simulation decision claim written?
- A simulation decision claim states what decision the result supports, what performance is observed, the conditions of observation, the intended population and the consequence of the decision. For example, a claim can specify whether a result supports practice feedback or a decision about independent operation. Writing the claim first prevents a scenario designed for development from being silently reused as high-stakes evidence.
- Does a more realistic simulator always make the result better?
- No. Fidelity is selected for the decision claim and the risk, not as a generic quality ranking. A simulator can be excellent for familiarisation or procedural practice without supporting regulatory credit or a broad workplace-readiness claim. The relevant question is whether its equipment, information flow, team conditions and consequences let the programme observe the particular decision it says the result represents.
- What makes a simulation score fair enough for progression or employment?
- A consequential simulation needs standardisation, prepared raters, published scoring rules and an appeal route. Standardisation controls the starting state, information, role-player responses, time window and allowable prompts. Raters calibrate on sample performances before high-stakes use. Scoring defines critical errors, partial credit, remediation and pass. Appeal preserves a documented basis for reviewing procedural defects.
- Why should a simulator configuration and rubric have versions?
- A consequential result needs a record of the actual scenario version, simulator configuration, administration conditions, rubric, scorer materials and incidents. A version identifier such as item 3.2 or rubric 2.1 makes later review possible; a label such as latest final revised does not. Without those records, an organisation cannot determine whether two scores reflect the same task or whether a changed setting altered the claim.
- Can a checklist turn a practice simulation into an assessment?
- No. A checklist can support observation during formative practice, but formative simulation does not become assessment merely because someone records actions. Assessment requires an explicit decision claim, an observation method, a scoring rule and an interpretation that the evidence can bear. When a result affects progression, credentialing or employment, standardisation, rater preparation, scoring controls and appeal also become necessary.
- What should happen when most experienced participants miss the same control?
- The programme investigates the scenario before treating the pattern as repeated individual failure. If 8 of 10 experienced participants cannot locate an unfamiliar control, the result can indicate a design or orientation defect. The control may not have been introduced as promised, or the simulated setting may not provide the psychological fidelity needed to interpret the action as workplace judgement.
- What belongs in a simulation debrief?
- A useful debrief reconstructs observed events, examines what worked and did not, identifies the reasoning and system conditions behind a gap, and agrees a change for future cases. Time-stamped facts support this work better than labels such as good leadership. A short team debrief can take about 3 minutes, while a learning debrief may need longer analysis of only the stated objectives.
- How can an organisation tell whether learning transferred from the scenario to work?
- An organisation checks transfer through later field observation and changed work conditions, not through scenario completion alone. It records the practice need, objectives, observations, corrective actions, owners and completion status, then observes whether the target behaviour occurs under operational conditions. Repeating a scenario with changed conditions can test adaptation, but it remains scenario evidence until workplace observation supports the transfer claim.
- Why are simulation logs risky input for an AI system?
- Simulation logs are structured, time-stamped records, so an AI system can treat them as more general evidence than they are. A log records behaviour in one scenario version, simulator configuration, prompt sequence and scoring model. If those boundaries are absent, the system can convert a scenario pass into an inferred job capability, recommend mobility or waive development without evidence from the real operating setting.
- What disclosures are required before people take part in a simulation?
- Participants need disclosures about purpose, data handling and safety boundaries. Purpose states whether the event is practice, assessment, research, a system test or a combination. Data handling states what recordings or notes exist, who can access them, how long they are held and whether they enter employment or education records. Safety boundaries identify prohibited actions and the way to pause.