🏆 BETT 2025 Kids’ Judge Award⭐ Tech&Learning Best of BETT 2026

AI-Proof Assessment: Grade the Process, Not the Product

Generative AI can write any lab report, but it can’t do a titration. WhimsyLabs grades what students actually do in the lab: their technique, their decisions, their safety habits. There’s nothing for AI to fake, and nothing for you to police.

94%of UK undergraduates use generative AI to help with assessed work[4]
40%how often AI and human graders agree exactly when marking the same essay[2]
82%of educators worry about students using generative AI on assignments[3]

Why Doesn’t AI Detection Work?

Direct answer: AI detectors can’t reliably tell human writing from machine writing, and their mistakes land on real students. Reported false-positive rates run between 5% and 20%[1], and a study presented at the 2024 American Educational Research Association conference found AI and human graders reach exact agreement only about 40% of the time[2].

Detection also turns the classroom adversarial. Students are treated as cheats until proven innocent, and the tools flag non-native English speakers and neurodivergent students most often. In Pearson’s 2025 formative-assessment research, 82% of educators said they were concerned about students using generative AI on assignments[3]. Two years of detection tools haven’t shifted that number, because catching AI text doesn’t bring back the learning the assignment was meant to produce. The fix isn’t better policing. It’s setting work AI can’t do.

What Is Process-Based Assessment?

Direct answer: Process-based assessment grades how a student works: their hypotheses, technique, decisions, and how they respond when something unexpected happens, rather than only the final answer they hand in. The reasoning and the physical actions have to be the student’s own, so there’s nothing a chatbot can produce on their behalf.

A correct molarity at the end of a titration tells you very little on its own. The student may have reasoned carefully, copied a neighbour, or guessed. The process tells you what the number hides: did they clear the air bubble from the burette, slow to dropwise near the endpoint, and repeat the anomalous reading? Pearson’s research found educators rank essays and multiple-choice questions as the formats most vulnerable to AI misuse, and simulations as the least[3]. Practical, simulated work is where assessment can still be trusted.

How Does WhimsyLabs Grade the Process?

Direct answer: WhimsyLabs logs every action a student takes in the virtual lab, from equipment handling to measurement precision, timing, and safety. Its AI then suggests grades across five dimensions (experimental procedure, data collection, calculations, lab safety, and scientific communication) for you to review and approve.

Each suggestion carries a confidence level, so you can see where the automated marking is dependable and where your judgement matters more. Student behaviour is compared with expert pathways for each experiment, which surfaces the gaps written reports never show, like the student who can explain why swirling matters but never swirls the flask. Follow-up questions are built from each student’s own experimental data, so a generic AI answer is no help.

Physical actions also remove the ambiguity that breaks text-based AI marking. A virtual pipette that recorded 2.47 mL transferred needs no interpretation. There are no synonyms for a physical action, and no wording to misread.

WhimsyLabs AI grading breakdown scoring experimental procedure, data collection, lab safety, and scientific communication for a student

Is the AI Safe for Students?

Direct answer: Yes, because WhimsyCat, our AI tutor and assessor, has no student chat window. It works everything out from students’ actions in the lab. Pupils never type prompts and never receive generated text.

The risks the Department for Education’s 2026 generative AI product safety standards target (manipulation, dependence, harmful generated content, jailbreaking) all start with an open chat channel pointed at a child[6]. WhimsyCat removes that channel by design. There is nothing to jailbreak, no conversation to grow dependent on, and no free text to filter. Signs of struggle or frustration, like repeated failed attempts or a long hesitation, are picked up from behaviour and flagged to the teacher, who stays the human in the loop.

Read the full breakdown in how WhimsyCat meets the DfE AI safety standards without a chat box.

WhimsyCat, the WhimsyLabs AI tutor, observing a student's virtual experiment without a chat window

Under the Hood: How Action-Based Grading Actually Works

Direct answer: Every session runs through five stages: capture, structure, assess, guide, and replay. Actions are logged in the simulation, compared against the expected pathway for that practical, and graded by a language model we fine-tuned for lab work. At no stage does a student type anything.

  1. Capture. Our action logger records every meaningful event in the lab across more than forty categories: reagent additions, pours and pipette transfers, heating and stirring, reaction milestones, equipment connections, and safety events. Hand and head movement is sampled five times a second, so hesitation and technique are part of the record. The logger also knows why things changed. It can tell a reagent the student added from one produced by a reaction, so a lucky accident never grades the same as a deliberate step.
  2. Structure. The raw stream is condensed into a compact session log: every action timestamped, with its safety level and outcome. The same log exports as a plain-English report, so nothing the AI sees is hidden from you.
  3. Assess. The log is compared with the expected actions and outcomes for that practical: what a competent scientist would have done, and what should have resulted. A language model we fine-tuned for this job suggests grades across the five skill areas, each with a confidence level, for you to review. It was trained on synthetic lab sessions we generate in-house, covering correct and flawed procedures at every skill level. It is never trained on pupils’ data.
  4. Guide. WhimsyCat runs off the same action stream. When the pattern says a student is stuck, it responds with guidance drawn from a library of pre-written, cached responses. Nothing is generated live in front of a student, and there is no text prompt anywhere in the student experience.
  5. Replay. Because the log is complete, any session can be replayed. You can scrub to the exact moment a titration went wrong and watch it happen, students can review their own technique, and every suggested grade has a visible evidence trail behind it.

Do Students Really Use AI That Much?

Direct answer: Yes. The Higher Education Policy Institute’s 2026 Student Generative AI Survey found 94% of UK undergraduates use generative AI to help with assessed work, and 12% put AI-generated text straight into their submissions[4].

That figure isn’t evidence of mass cheating. It’s evidence that grading things AI produces effortlessly no longer measures learning. While the mark rewards the final document, using AI to make it is the rational move. Grade the process instead and the incentives flip: students can still use AI to revise, but the skills being marked (technique, observation, adapting when a practical goes wrong) have to be shown first-hand.

What Do the OECD and DfE Say?

Direct answer: Both point the same way. The OECD recommends process-oriented assessment for the AI age, and the DfE’s AI safety standards reward tools that avoid open generative chat with pupils.

  • The OECD Digital Education Outlook 2026 concludes that assessment focused only on final outputs is becoming inadequate, and urges educators to look at how students engage with learning. It also warns about cognitive offloading, citing studies where AI-assisted students completed tasks 48% more successfully but performed 17% worse once the AI was taken away[5].
  • The DfE’s generative AI product safety standards (updated January 2026) require AI tools in schools to guard against manipulation, dependence, harmful content, and training on pupils’ work without consent. WhimsyCat meets these requirements by architecture, not by retrofitted filters[6].
  • Pearson’s Assessment Evolved report recommends using AI as a force multiplier for formative assessment while educators keep control[3], and the HEPI 2026 survey concludes institutions must make sure AI enhances learning rather than diminishing it[4].

Go Deeper: The Research Behind This Page

Make Your Practical Assessment AI-Proof

See how process-based grading works in your own lessons, from the student’s lab bench to your grading queue.