Rolling out a criterion checker to your department
Introducing an AI checker to a department is a change-management exercise wearing a technology costume. Done wrong, it becomes "that thing we bought and nobody uses". Done right, it quietly becomes how the department marks. The rollout that works.
Step 1 — Frame it before you buy anything
The framing that lands with teachers: "the tool reads; you judge." Advisory verdicts with cited evidence, teacher decides everything, no grade emission — if a tool can't say that sentence truthfully, it's the wrong tool (the vendor questions). Say the frame in the first meeting; keep saying it in every demo.
Step 2 — Pilot with the willing, on one unit
- One unit, one round, two teachers. Small enough to abandon, real enough to trust.
- Run the comparison that matters: the teachers mark their normal way and run the checker; compare verdicts per criterion. Agreements build trust; disagreements build calibration (consistency guide).
- Measure honestly: marking time per script, record completeness, and — the real metric — how many Pass-criterion gaps the checker surfaced that manual marking would have caught at IV instead.
Step 3 — Tell your verifier before your verifier asks
The compliance conversation is short if you initiate it: "We use an advisory criterion-checking tool. It reads evidence against the specification's criteria and produces per-criterion verdicts with cited evidence. Teachers review every verdict and make every decision. Here's the sample output and our workflow." Verifiers care about who decides and what the records show — both of which your pilot already demonstrates (the SV guide for the records side).
Step 4 — Expand with the sceptics, not around them
The sceptical teacher is your best pilot for round two: they'll find every false verdict, and every one they find that gets one-click overridden is a demonstration of the teacher-decides boundary. Tools win departments over through demonstrated override-ability, not evangelism.
Step 5 — Bake it into the workflow, not the culture
- Standardisation meetings use the checker's shared verdicts as the comparison grid — the method.
- Submission-time checks become the department's default — prevention economics.
- IV sampling starts from the checker's verdict sheet — the IV guide's triage-first pattern.
The rollout metrics that prove it worked
| Metric | Watch for |
|---|---|
| Marking time per script | Falling, mostly from the locating pass |
| IV disagreement rate | Falling — records carry citations |
| Resubmission volume | Falling — mechanical Pass gaps caught at submission |
| Record completeness at SV time | 100% without a June scramble |
| Tool-vs-teacher verdict agreement | Rising over rounds — calibration, not coincidence |
FAQ
What if only half the department adopts it?
That's fine at first — but standardisation needs shared tools. Review after a term: if the adopters' records look different from the non-adopters', that difference is the argument.
Does using a checker change what verifiers sample?
Your records get stronger, so if anything sampling gets easier. The verifier's questions are the same ones your teachers already answer with cited evidence — see the SV guide.