Home / Blog / Feedback & workload
Feedback & workload

Getting marking consistency across assessors

Consistency isn't a personality trait — it's a system. When two assessors in one BTEC department produce different verdicts for the same evidence, the cause is almost never values; it's that they never wrote down a shared threshold. The fix has four parts, and they're all cheap.

Part 1 — Write the thresholds together

Before each round, the team writes one line per criterion: "met means X in this brief." Ten minutes per unit. The document is the department's standard made visible — and it's what new assessors inherit instead of folklore. This is the same habit that makes solo marking stable (verdict vocabulary), scaled up.

Part 2 — Standardise on real scripts

Once per round, the whole team marks the same (anonymised) script independently, then compares criterion by criterion, not script by script. The disagreements are gold: each one marks a threshold that isn't shared yet. Resolve, update the threshold document, move on. The Lead IV owns this meeting — it's the highest-leverage hour in their calendar (the Lead IV role).

Part 3 — Make comparison mechanical

The fastest route to alignment is putting the same evidence in front of everyone with the verdicts already visible. When every assessor sees the same per-criterion result with cited evidence, the conversation shifts from "I feel it's a Merit" to "look at A.M2's evidence". Departments run exactly this with checkb.tech: one shared script, the by-criterion view on screen, five minutes of genuine comparison instead of an hour of assertion. And because the same evidence always yields the same verdicts — deterministic checking — there's no "the tool marked differently on Tuesday" to argue about.

Part 4 — Check for drift

Consistency decays silently. Three cheap detectors:

  • Verdict distribution by assessor. If one assessor's "met" rate for a criterion is far from the team's, that's a conversation — cohort analytics shows this per criterion across a class in seconds.
  • Re-mark three early scripts at the end of the round. If your own verdicts moved, you drifted; the middle of the batch needs a second look.
  • IV disagreement rate. A few disagreements per round is the IV system working. Zero disagreements for a term means nobody's sampling honestly; a spike means a threshold is contested.

Why consistency is a compliance issue, not a nicety

Standards Verifiers exist to compare your centre's standard against the national one; internal inconsistency is the fastest way to invite deep sampling. A department that can show a thresholds document, standardisation minutes and drift checks is presenting a system, not a set of individuals. That documentation posture is most of what the SV preparation guide recommends.

FAQ

How often should standardisation run?

Per assessment round — not per year. The evidence changes with each brief; the shared threshold has to be re-anchored to it.

What if we're a one-person department?

Cross-standardise with a neighbouring subject's assessor or partner centres, and lean harder on determinism: a tool whose verdicts don't drift becomes your second opinion.

Check the criteria before the IV does.

checkb.tech reads learner evidence against the official Pearson criteria and reports every criterion as met, partly met or not met — with the evidence cited. It never awards a grade; the teacher stays the assessor.

Try checkb.tech free →