Getting marking consistency across assessors
Consistency isn't a personality trait — it's a system. When two assessors in one BTEC department produce different verdicts for the same evidence, the cause is almost never values; it's that they never wrote down a shared threshold. The fix has four parts, and they're all cheap.
Part 1 — Write the thresholds together
Before each round, the team writes one line per criterion: "met means X in this brief." Ten minutes per unit. The document is the department's standard made visible — and it's what new assessors inherit instead of folklore. This is the same habit that makes solo marking stable (verdict vocabulary), scaled up.
Part 2 — Standardise on real scripts
Once per round, the whole team marks the same (anonymised) script independently, then compares criterion by criterion, not script by script. The disagreements are gold: each one marks a threshold that isn't shared yet. Resolve, update the threshold document, move on. The Lead IV owns this meeting — it's the highest-leverage hour in their calendar (the Lead IV role).
Part 3 — Make comparison mechanical
The fastest route to alignment is putting the same evidence in front of everyone with the verdicts already visible. When every assessor sees the same per-criterion result with cited evidence, the conversation shifts from "I feel it's a Merit" to "look at A.M2's evidence". Departments run exactly this with checkb.tech: one shared script, the by-criterion view on screen, five minutes of genuine comparison instead of an hour of assertion. And because the same evidence always yields the same verdicts — deterministic checking — there's no "the tool marked differently on Tuesday" to argue about.
Part 4 — Check for drift
Consistency decays silently. Three cheap detectors:
- Verdict distribution by assessor. If one assessor's "met" rate for a criterion is far from the team's, that's a conversation — cohort analytics shows this per criterion across a class in seconds.
- Re-mark three early scripts at the end of the round. If your own verdicts moved, you drifted; the middle of the batch needs a second look.
- IV disagreement rate. A few disagreements per round is the IV system working. Zero disagreements for a term means nobody's sampling honestly; a spike means a threshold is contested.
Why consistency is a compliance issue, not a nicety
Standards Verifiers exist to compare your centre's standard against the national one; internal inconsistency is the fastest way to invite deep sampling. A department that can show a thresholds document, standardisation minutes and drift checks is presenting a system, not a set of individuals. That documentation posture is most of what the SV preparation guide recommends.
FAQ
How often should standardisation run?
Per assessment round — not per year. The evidence changes with each brief; the shared threshold has to be re-anchored to it.
What if we're a one-person department?
Cross-standardise with a neighbouring subject's assessor or partner centres, and lean harder on determinism: a tool whose verdicts don't drift becomes your second opinion.