Choosing a BTEC checking tool: the 12-point checklist
You've decided a checking tool might earn a place in the workflow. The market is loud and the demos are polished. This is the evaluation checklist — print it, score every tool against it, watch the answers separate the tools that help from the tools that generate findings.
Assessment fidelity (the non-negotiables)
- Specification pinning. Can the tool name the exact spec issue and units it covers? "BTEC-aligned" is not an answer; "International L3 Business Issue 4, 43 units" is.
- Criteria verbatim. Ask to see one unit's criteria as the tool holds them, against the spec's page. Paraphrased criteria mean paraphrased standards.
- Brief version handling. V1/V2 (and set choices) change the scenario and data your evidence gets read against — why it matters. How does the tool know which brief a batch used?
- Verb-level checking. Show me a described-not-analysed passage against an analyse criterion. Does the verdict and comment name the verb gap?
Verdict quality (the trust tests)
- Cited evidence on every verdict. No citation, no trust — you'd be confirming blind (verdict vocabulary).
- Determinism. Same submission, same brief, same config — same verdict, this month and next? Non-deterministic tools make appeals unwinnable.
- Override design. When you disagree: one click, your decision recorded? Or a support ticket?
- The grade boundary. Nothing in the product emits grades to learners or presents itself as the decision-maker? This is the compliance keystone — the full vendor questions.
Workflow fit (the daily-use tests)
- Exports for verification. Produce a per-learner, per-criterion record with evidence locations — the shape verifiers sample for (records guide). A score PDF is not it.
- Batch reality. How does 30 scripts × 3 file formats actually go in? Drag-and-drop on a Tuesday, or per-file forms that take longer than marking?
Safety and commercial (the boring tests that matter most)
- Data protection. Written no-training-on-learner-data confirmation, UK/EU GDPR posture, retention control, tenant isolation — our own answers are public; insist every vendor's are.
- Pricing vs your volumes. Scripts per month you actually mark × price per tier, against the marking hours saved — a measured marking rate makes this concrete. Free tiers that fit a real batch are for testing exactly this.
How to run the evaluation
One unit, one real batch, two teachers, verdicts compared against manual marking — the pilot design from the department rollout guide. Score the checklist during the pilot, not from the demo: demos are choreographed, batches are honest.
FAQ
What weight matters most?
Fidelity (1–4). A fast, cheap tool reading against the wrong criteria is confidently wrong at scale — the worst possible product category.
Should the tool also detect AI writing?
Screening for authenticity is valuable, but it's a different function from criteria checking; evaluate both functions separately, and never let a detector score masquerade as an authenticity decision (the honest guide).