Assessment proxy
Instructional Design Cannot Repair a Bad Capability Claim at the End
A polished course can be expertly designed and still prove the wrong thing. The capability claim, conditions and assessment evidence have to be decided before the catalogue, media and completion report arrive.
What the wrapper says
Instructional design selects objectives and learning experiences, while assessment tests whether the resulting evidence can support the claimed capability decision.
The all-clear
Instructional design and learner feedback make learning efficient, accessible and coherent when they are not mistaken for evidence of achieved competence.
What is durable
The defined performance a learner can demonstrate under the conditions that matter for the intended decision.
Your course has a cinematic opening.
Accessible captions.
Tasteful motion graphics.
And an exit survey with enough agreement emojis to power a quarterly business review.
It also has a completion record.
Somebody now wants to put “capable of incident triage” into the skills system.
This is where the course starts being asked to do something it was never designed to do.
Instructional design is not a capability claim.
It is the disciplined work of selecting objectives, activities, media and learning conditions.
Assessment is the disciplined work of asking whether a result can support a particular decision about performance.
A completion record is an institutional wrapper around participation in the experience.
Three related objects.
And the LMS button labelled Award skill does not know the difference between them.
Start with the decision that needs evidence
The first sentence is not “the module should be engaging.”
It is this.
What will this result let somebody decide, and under what conditions was it produced?
Then name the population, the decision maker, the consequence, and whether the decision can be reversed.
A short diagnostic at entry can identify missing prerequisites.
It is not secretly a grade, a promotion signal, or permission for independent operation, merely because it returned a number.
A more consequential decision needs a correspondingly credible chain of evidence.
That sentence gives the objective a job.
In Dick and Carey’s ten-step systems approach, the team writes performance objectives with a performance, a condition and a criterion.
Then it develops criterion-referenced assessment instruments.
Then, and only then, it develops instructional materials.
The assessment is step five.
The materials are step seven.
This is occasionally received as an administrative inconvenience by the team holding a nearly complete storyboard.
It is actually the moment the programme decides what it is claiming.
ADDIE says the same thing in another dialect.
Analysis, Design, Development, Implementation, Evaluation.
Task inventories become measurable terminal and enabling objectives during Design, before the media production machinery starts humming.
The military IPISD specification makes that machinery unusually explicit, with 19 standardized sub-tasks and formal review gates.
The gates are not there because instructional designers enjoy paperwork.
Nobody enjoys paperwork.
They are there because a course cannot acquire a missing criterion by becoming more beautifully narrated.
A good objective can still be attached to the wrong evidence
The usual failure is tidier than a bad course.
The outcome says analyse.
The final check asks candidates to recognise a definition.
Both concern the same topic.
The evidence does not concern the same performance.
That is not alignment.
That is a shared noun wearing a lanyard.
An assessment blueprint makes the mismatch visible before your reporting system makes it permanent, which it will.
Content or task areas down one side, performance demands across the other: recall, interpretation, procedure, explanation, judgment, production.
Your weighting follows the decision, not the calendar.
Equal treatment time is not a defensible rationale.
Neither is the fact that the instructional team has already made several videos about the subject.
Format follows evidence too.
A one-best-answer item can compare judgment efficiently across a large group.
A short constructed response can show a chain of reasoning.
A performance station can show execution under stated conditions.
A portfolio can show work over time.
Oral questioning can show explanation, provided the prompt and the rubric do not drift with whoever is asking.
The format is not the capability.
It is a way of eliciting evidence about the capability.
So when a client handover requires observed performance or an artefact, a multiple-choice check does not become more valid through elaborate distractors, a timer, and a progress bar reading “Almost there.”
It is still evidence of what the item elicited.
The rubric is part of the claim
Tasks arrive with their scoring rules, an intended answer or exemplar, a rationale, and a subject-matter check.
For performance work, you define the observable dimensions, the evidence, the levels and the non-compensable critical errors before anyone begins.
Levels have to name visible features.
“Good” and “excellent” are not level descriptors.
They are wishes with formatting.
The distinction matters most where a total score can hide a decision you would not actually make.
A safety-critical criterion is conjunctive: it has to be met, and no amount of good work elsewhere converts a critical omission into an acceptable result by arithmetic.
The same discipline applies to the reported pass.
A claim states the routine, the critical criterion, the condition and the rule.
Under these observed conditions, this result supports a decision that the trainee can complete this task with no critical omission.
“Completed Incident Triage Foundations” identifies a catalogue object.
It is not the same kind of statement.
The design work is genuinely good, which is the problem
None of this argues against polished courses, learning objectives, learner feedback, or instructional designers who can make a difficult procedure comprehensible before lunch.
Instructional design solves real problems.
It makes learning coherent.
It sequences practice.
It catches accessibility barriers before they become an accidental test of keyboard speed or platform familiarity.
It gives your dispersed workforce a repeatable experience, rather than a collection of charmingly contradictory manager explanations.
For complex work, the 4C/ID model is especially clear about why content alone is insufficient.
It combines four components: authentic whole learning tasks, supportive information for the non-routine aspects, just-in-time procedural information, and part-task practice for routines that need automaticity.
The learner meets whole tasks of increasing difficulty, and the scaffolding fades toward autonomous performance.
That is excellent instructional design.
It does not mean that viewing the sequence proves autonomous performance at work. It tells the team how to build toward it.
Agile models are useful for a different reason.
A savvy start produces low-fidelity functional prototypes in a day or two, then iterative loops move through a design proof to alpha, beta and gold.
Fast feedback finds a confusing interaction before a large release.
It cannot turn an unspecified authorization decision into a valid one.
The design deserves credit for the thing it is evidence of: a considered learning experience.
The mistake begins when that evidence gets promoted into proof that a different thing happened.
Pilot the response process, not just the slide deck
The course preview goes very well.
The expert approves the examples.
The programme manager enjoys the animations.
The authoring tool produces a completion badge with a reassuring amount of blue.
None of that tells you what your candidates think the task requires.
Dick and Carey’s formative sequence has three stages for exactly this reason.
One-to-one work with three to five representative learners.
Small-group trials with eight to twenty.
Then a field trial with thirty or more, in the real operating environment.
The small-group stage checks whether learners can execute independently, and whether the item-level evidence supports the intended mastery threshold without external coaching.
At field trial, server bandwidth, instructor fidelity, administrative handoffs and the live LMS all become part of the evidence.
A beautiful assessment that behaves differently when the VPN has opinions is measuring more than the intended capability.
In a thirty-minute pilot, three candidates stopping at the same instruction are evidence about the instruction.
They are not a small rebellion against the course.
So your pilot records completion time, clarification requests, navigation errors, resource use and interruptions.
Then the team can compare the intended response process with the one actually produced.
Then lock it.
Version, instructions, permitted resources, accommodation process, scorer materials, incident process.
Item 3.2 and rubric 2.1 make a later review possible.
“Latest final revised” makes a later review into archaeological fieldwork.
Satisfaction is a useful report about satisfaction
The post-course reaction score is often the cleanest number on your dashboard.
Understandably.
It measures reaction: satisfaction and perceived utility.
Nearly every programme collects it.
It is useful.
A course your learners cannot navigate, understand or tolerate has a real design problem.
It is not a report of workplace performance.
The next level up measures knowledge and skill acquisition, through pre- and post-tests or rubric-scored simulation.
Roughly three-quarters of programmes get that far.
The level after that measures on-the-job application, through supervisor audits at 30, 60 and 90 days, or peer observation.
Implementation falls to something like a fifth of programmes.
The last level measures organisational results.
Defect rates.
Sales growth.
Retention.
Fewer than one programme in ten reaches it.
That drop is not a moral indictment of the people running learning.
Observing workplace behaviour and results costs more than sending a survey link.
It is a reminder that those reports make distinct claims.
A 95 percent pass rate on the learning assessment does not contradict a later finding that fewer than a fifth of graduates apply the procedure at 60 days.
It tells you to investigate management reinforcement and workplace friction, rather than congratulating the completion report for solving a different problem.
The platform deserves separate inspection too.
The System Usability Scale runs ten items, five positive and five negative, and produces a composite out of 100.
Its baseline average is 68.
A score above 80 sits in the top tenth.
Interface friction can make an assessment partly about navigation, when navigation was not in the claim.
The learner may have enjoyed the course.
The system may have been usable.
The learner may also be unable to perform the task under the conditions that matter.
All three can be true at once, which is inconvenient for a single-field report.
Abundant content, unchanged claim
Generative production makes content abundant.
A team can now produce a module, a voiceover, a knowledge check and a completion signal at a speed that would previously have required a request form, a vendor, and somebody discovering the brand-font file was missing.
Which makes the distinction more urgent, not less.
An AI system will happily ingest course titles, objective labels, quiz scores, completion dates and learner reactions.
They are structured.
They look like workforce data.
It can infer a skills record, recommend an assignment, or flag a person for promotion, with no visible malice and a beautifully formatted confidence score.
But a label is not evidence.
A completion is not performance.
And a satisfaction score is not authorization.
So your record has to retain the decision claim, the condition, the task format, the scoring rule, the administration version, the result and the limitation.
A qualified result against a stated operational threshold, under one standard administration, is a qualified result.
“Competent” is a much larger claim, and your database should be required to explain what it means by it.
The durable object is the defined performance a learner can demonstrate, under the conditions that matter for the decision.
The course is the learning experience.
The completion record is a participation wrapper.
The objective label is a design wrapper.
The score is evidence only to the extent that its task, conditions and interpretation support the decision.
Build excellent learning experiences.
Measure their usability.
Ask learners what worked.
Then decide, before the catalogue arrives, what has to be true before you write a capability claim into a system that will remember it for years.
What the record establishes
- Write an objective with a condition, action, and criterion tied to a named decision.
- Build an assessment blueprint from the performance claim before writing course media.
- Choose response formats that elicit the evidence needed for the decision rather than the cheapest score.
- Draft tasks, scoring rules, and critical-error criteria together.
- Pilot instructions, technology, and response processes before a consequential release.
- Separate reaction, learning, workplace behavior, and organizational-results reports.
- Report performance results with their conditions, decision rule, and limitations rather than as a completion label.
Asked in the review
- What does a course completion record actually prove?
- A course completion record proves that a learner completed the defined course activity under the platform's recorded conditions. It does not, by itself, prove that the learner can perform a job task, explain a failure, or apply a procedure at work. A capability decision needs evidence from a task that elicits the claimed performance, a stated scoring rule, and a record of the administration conditions.
- What has to be decided before anyone starts building the course?
- The organisation states the decision claim before selecting a format or authoring media: using the result, it decides whether a learner can perform a named action under named conditions. The claim identifies the performance, population, consequence, and decision maker. A 15-minute entry diagnostic can identify missing prerequisites, but it does not become evidence for a grade or an independent-operation decision merely because the score is available.
- How does an objective become something that can be assessed?
- A performance objective states the action, the conditions, and the criterion. The action describes what the learner does; the condition states resources, environment, or constraints; and the criterion states the standard for acceptable performance. Dick and Carey's systems approach places performance objectives before criterion-referenced assessment instruments, so the assessment can be designed to gather evidence for the objective rather than retrofitted to a finished course.
- Why is a learning objective label not enough for a skills record?
- A label such as analyse names a desired demand but does not show that a task actually elicited analysis. A recognition item can share the topic and use the same verb while measuring recall knowledge instead. A skills record needs the task demand, conditions, scoring evidence, and decision rule that support the stated capability. The label is useful design metadata; it is not proof of achieved performance.
- How should a course team choose an assessment format?
- The format follows the evidence required by the decision. One-best-answer items efficiently compare judgment across many candidates. Short constructed responses reveal a chain of reasoning. Performance stations show execution under conditions, portfolios show work over time, and oral questioning can test explanation when prompts and rubrics are stable. Fast scoring is a cost consideration, not evidence that every objective is appropriately measured by multiple choice.
- What does an assessment blueprint do that a course outline does not?
- An assessment blueprint maps content or task areas to performance demands such as recall, interpretation, procedure, explanation, judgment, or production, and assigns planned points, tasks, or time to each cell. A 40-point assessment can allocate 16 points to routine procedures, 12 to diagnosis, 8 to safety decisions, and 4 to documentation when that distribution reflects the decision claim. A course outline does not establish that coverage.
- When do tasks and scoring rules need to be written?
- Tasks and scoring rules are drafted as one production unit before candidates respond. Each task has an intended answer or exemplar, rationale, source or subject-matter check, and scoring rule. For performance work, the rubric defines observable dimensions, evidence, levels, and any non-compensable critical errors before administration. This prevents a scorer from inventing a criterion after seeing a response or converting an attractive outcome into a different claim.
- What should a pilot test besides whether people like the course?
- A pilot tests the instructions, technology, response process, timing, resources, and scoring process under conditions that resemble delivery. Dick and Carey's formative sequence uses one-to-one reviews with 3 to 5 learners, small-group trials with 8 to 20 learners, and field trials with 30 or more learners. In a 30-minute pilot, repeated stops at the same instruction are evidence that the instruction needs repair.
- Can a high satisfaction score show that people are competent?
- No. Kirkpatrick Level 1 measures reaction: participant satisfaction and perceived utility. Level 2 measures learning through instruments such as tests or rubric-scored simulations, Level 3 measures workplace behavior, and Level 4 measures organizational results. The levels answer different questions. A favorable reaction report can support a claim about learner experience, but it does not establish application on the job or an authorization decision.
- Why can a polished learning platform still produce bad evidence?
- A learning platform can be visually coherent and still add interface friction that changes what an assessment measures. The System Usability Scale uses 10 items, five positive and five negative, to produce a 0-to-100 usability score. A score below 68 is associated with severe friction, while 80.3 or higher is the Grade A benchmark. Navigation difficulty can become an unannounced demand in the assessment result.
- What should appear in a report that says someone passed?
- A performance report states the observed score, the applicable category rule, the conditions of administration, and material limitations. For example, 19 of 24 can be reported as meeting an operational threshold of 18 under one standard administration. That is more precise than competent. A one-point margin prompts review of scoring, incidents, and retake or appeal policy because the category boundary carries the decision.
- What changes when generative AI can make a course almost immediately?
- Generative production makes course pages, narration, checks, and completion records easier to create at scale. It does not create the evidence needed for a capability claim. An automated system should retain the objective, required conditions, task format, scoring rule, administration record, and limitations with any performance result. Without those fields, an AI recommendation treats a polished completion as demonstrated competence and can make an unsupported assignment or promotion decision.