Pacific Technology SolutionsTraining & software for the motor industry

Bay & Classroom / Training Design

02Training Design

Assessments That Predict Whether a Technician Can Do the Job

A multiple-choice test measures whether someone can pass a multiple-choice test. What predicts workshop performance, and how to build it into courseware. For a practical discussion of manipulated activity signals, see this page.

Most technical courseware ends in a multiple-choice test with a pass mark. It is cheap, it is automatic, it produces a number, and it measures recognition of correct statements.

Whether the technician can perform the work is a different question, and the gap between the two is why a network can report high pass rates alongside unchanged comeback rates.

Why recall tests do not predict performance

Recognition is easier than recall, and recall is easier than performance. A technician who can identify the correct answer among four cannot necessarily produce it from nothing, and someone who can produce it cannot necessarily execute it at a vehicle under time pressure.

The distractors do the work. Multiple-choice difficulty is largely a property of how plausible the wrong answers are, which means the test measures item-writing quality as much as knowledge.

The context is absent. Diagnosis is a sequence of decisions made with incomplete information. A question with all the information present and one decision to make is not that task.

And test-taking is a skill of its own, which transfers between tests and not to the workshop.

What predicts performance better

In rough order of predictive value against cost.

Doing the actual task, observed. A structured practical assessment at a vehicle, scored against defined criteria. Expensive, and nothing else predicts as well.

A scenario worked through from symptom to conclusion. Present a complaint, let the technician request information — freeze frame data, a measurement, a test result — and evaluate the path as well as the answer. This is much closer to the job than any question format and it can be delivered as courseware.

Interpretation of real artefacts. A waveform, a live data screen, a fault code set, a photograph of a component. "What does this tell you and what would you check next" is a strong item type and it is cheap once the artefacts exist.

Recall rather than recognition. Asking for the specification rather than offering four.

Ordering and sequencing. Put the diagnostic steps in the correct order, which cannot be guessed as easily.

And the weakest that is still worth having: well-written multiple choice on knowledge that genuinely is recall — torque figures, service intervals, safety requirements.

Designing an assessment that reflects the job

Start from the task, not from the content. What must the technician be able to do at the end? Write the assessment first, then build the course toward it. Building the course and then writing questions about it produces questions about the course.

Sample the job, not the syllabus. If 80% of the work is three procedures, the assessment should weight them accordingly rather than covering every module equally.

Include the conditions. Time pressure, incomplete information, an intermittent fault. A technician who performs when everything is available may not when it is not.

Include the decision not to proceed. Knowing when to escalate, when to stop, and when the information is insufficient is part of competence and it is almost never assessed.

Set the pass mark from the consequence. Safety-critical work has a different threshold from routine maintenance, and a single network-wide pass mark is a decision made by not deciding.

Practical assessment without a classroom

The obstacle to observed assessment is cost, and there are intermediate options.

Structured workplace assessment. A supervisor or senior technician scores the technician performing real work against a defined checklist. Cheap in materials, expensive in senior technician time, and it needs the assessors trained or the scoring drifts.

Photograph or video evidence submitted by the technician of a completed task, assessed remotely. Increasingly practical and it needs clear criteria to be consistent.

Simulation, where the fault can be reproduced. Best for rare or dangerous conditions. See when simulation beats hands-on.

Record it with xAPI, which is designed to capture activity outside a browser session and is what makes workplace assessment aggregable at network level. See standards.

What to avoid

Unlimited retakes with the same questions. The technician eventually passes by elimination and learns nothing. Randomise from a pool, or limit attempts, or both.

Questions answerable from general knowledge without the course.

Trick items. They measure carefulness, not competence, and they damage credibility.

Assessing what is easy to assess rather than what matters. Torque specifications are easy; diagnostic reasoning matters more.

A single test at the end. Spacing assessment through a course produces better retention and identifies problems while they are correctable.

And treating completion as the record. A completion is a record that someone reached the end. See measuring training effectiveness.

Validating that it works

The step that closes the loop and is almost never taken.

Compare assessment scores against workshop outcomes — comeback rate, efficiency on the covered operations, escalations — for the same technicians. If they do not relate at all, the assessment is not measuring the thing.

Check item performance. Questions that everyone passes and questions that everyone fails carry no information. Questions that the strongest technicians get wrong are usually badly written rather than difficult.

Ask experienced technicians to sit it. If people who demonstrably do the job well fail the assessment, the assessment is wrong. This costs an hour and it is the fastest validity check available.

Review after the first cohort, before the network-wide rollout.

For a manufacturer's training function

A network reporting pass rates has no capability data. Two dealers at 95% pass with very different comeback rates is a fact about the assessment.

Assessment design is where the leverage is. Course content can be excellent and the network still cannot tell who can do the work.

Standardise the criteria, not just the questions, if workplace assessment is used. Two assessors scoring the same performance differently makes the data worthless.

And publish what the assessment is for. Technicians who believe it is a formality treat it as one; those who believe it identifies genuine capability treat it differently.

The short version

Multiple choice measures recognition. Performance is a different construct and the two can move independently.

Scenario items, artefact interpretation and ordering are better and are still deliverable as courseware.

Observed practical assessment predicts best and costs senior technician time — which is why the intermediate options matter.

Write the assessment before the course, from the task rather than the syllabus.

And have experienced technicians sit it. If good technicians fail, the assessment is wrong, and that takes an hour to find out.

For public guidance on designing and evaluating training, consult CDC training-development resources.