12 May 2026 · 9 min read
How to tell if students used ChatGPT: a fair approach

If you set written homework, you've probably faced it: a submission that reads too polished, too structured, with none of the rough edges of real student thinking. Your gut says something's off. The question is what to do about it - and that decision matters more than most guidance acknowledges.
The first instinct for most teachers is to paste the work into a free AI detector. That approach has a well-documented accuracy problem, and in some cases it has harmed students who wrote their work honestly. There's a better way to approach this.
Why text-based AI detectors aren't the answer
Tools that score how 'AI-like' prose reads work by measuring statistical predictability - how expected each word choice is relative to what a language model would produce. The logic is reasonable. The results aren't reliable enough for educational decisions.
A 2023 Stanford study tested seven widely used detectors and found they flagged non-native English essays as AI-generated at rates up to 61%. The same tools had much lower false positive rates for native speakers doing identical assignments. ESL and EFL learners use more predictable vocabulary and conservative sentence structures not because they're using AI, but because they're writing carefully in a second language - and that pattern looks identical to AI output to these tools.
Several universities in the UK and US have quietly pulled AI detectors from their academic misconduct processes after receiving complaints from international students who were wrongly flagged. Turnitin, one of the most widely used tools, added AI detection then had to walk back confidence levels after documented false positive issues. Vanderbilt University disabled its AI detection tool entirely.
A false positive on an honest student's essay isn't a minor inconvenience. In formal misconduct proceedings it can mean a failed mark, an official academic record, or for international students, implications for visa status. These stakes mean detection needs a higher standard of evidence than current text-based tools can reliably provide.
What two categories of signal actually exist
There are two fundamentally different types of detection signal, and understanding this distinction changes how you approach the problem.
| Signal type | What it measures | Reliability | Fairness concern |
|---|---|---|---|
| Output signals | How the finished prose reads - style, vocabulary, structure | Low and declining as models improve | High: flags ESL writers at elevated rates |
| Process signals | How the work was entered - paste events, typing rhythm, session time | Higher: harder to fake, language-neutral | Low: identical for all genuine writers regardless of language |
What process signals actually reveal
When work is submitted through a tool that captures writing behaviour, the session record tells you things the finished text cannot. Genuine writing has a recognisable rhythm - typing happens in bursts with natural pauses, content grows gradually, edits appear throughout. A 500-word essay typically takes 35-60 minutes for a student who's genuinely thinking.
AI-assisted shortcuts look different in the session data. The most common patterns:
- A single large paste event accounting for most of the final word count, often appearing early in the session
- Session duration of 3-8 minutes for a 500+ word submission (implausible for genuine composition)
- Typing that shows an unusually even cadence - consistent with transcribing pre-written text, not composing
- Almost no deletion or revision activity, even in a lengthy submission
- The document reaching near-final length without any visible growth prior to a paste event
How to collect process evidence
If you're working with submissions you already have - emailed documents, pasted text - you've lost the most useful evidence. The writing history isn't embedded in the file.
The practical solution is to collect written work through a tool that captures session behaviour as students write. Learnaway records the timeline of writing events - when typing happened, paste event sizes, window focus changes, session duration - without recording the actual characters typed. Teachers see a process fingerprint, not a transcript. Tools like Google Docs also have version history, though it's less granular.
Setting up process capture before the assignment removes the need to react after the fact. You have the evidence whether you need it or not, and students know it.
Making it a conversation, not an accusation
Whatever evidence you have, a conversation is almost always more revealing than a formal accusation. Lead with curiosity rather than suspicion. The most useful question is also the most neutral: 'Can you walk me through how you approached this?'
A student who wrote the work will have a process to describe: where they started, what was difficult, what they changed. Students who submitted AI-generated work often struggle with this question because there's no process to recall - they didn't make the choices in the text.
If you have specific process data, you can raise it concretely without accusation: 'I noticed a large block of text appeared at minute three - can you tell me more about that part of your process?' This opens a door. A student who drafted in a separate app and pasted it in has a completely legitimate answer.
When the signals add up to something
No single signal is sufficient grounds for formal action. Together, three things make a stronger case: anomalous session data, inconsistency with the student's previous work quality, and an inability to explain specific choices or parts of the submission.
If all three are present, you have reasonable grounds for a formal referral. Document what you have first - a timestamped process log is substantially stronger evidence in misconduct proceedings than a prose score from a detector. One is a record of what happened; the other is a probabilistic claim about statistical patterns.
If only one or two signals point to concern, use them as the basis for a supportive conversation rather than an escalation. Most of the harm in these situations comes from acting before there's enough to go on.
Try Learnaway with your next homework
Related articles
How to talk to a student you suspect used AI on an assignmentAccusations can backfire badly. Here's how to turn AI suspicion into a fair, evidence-based conversation - and when to leave it at that.
Why AI detectors flag ESL students as cheaters - and how teachers can avoid itText-based AI detectors are systematically biased against non-native English writers. Here's the research, the legal risk, and a fairer detection approach.
How to detect AI writing in student work: a practical guideText-based AI detectors are unreliable and unfair to ESL writers. A better approach examines how work was written, not what it says - here's a method that holds up.