By the Deslop desk · 17 Aug 2026 · Sources linked at the end
You wrote it. You remember writing it. You remember the afternoon you spent on the second paragraph.
And now there is an email in your inbox with a percentage in it, and the percentage is high, and nobody seems interested in the fact that you were there.
The first thing worth knowing is that this happens constantly, it is well documented, and the tools doing it are far less certain than the email suggests. The second is that the number in that email cannot prove what your institution thinks it proves. Not because detectors are badly made, but because of what they measure.
A detector cannot see who typed. It has no access to your afternoon. What it does is measure how surprising each word is, given the words before it, and then decide whether the text as a whole is too unsurprising to be human.
That is the whole mechanism. Predictability in, verdict out.
The problem is that plenty of human writing is highly predictable, and not because the writer did anything wrong. Clear, plain prose is predictable. Conventional structure is predictable. Writing in a language you learned as an adult is predictable, because your vocabulary is narrower and your sentence patterns are more regular. So is writing to a formula you were taught, which is what a five-paragraph essay is.
There is one experiment that settles this, and it is the most useful thing here.
Read the bottom half again. American schoolchildren wrote those essays. Simplifying their word choices took the AI flag rate from about 5% to about 57%.
So the tools are not measuring who wrote the text. They are measuring vocabulary sophistication, and reporting the answer as though it were authorship. If you write plainly, or if English is your second language, or if you were taught to write in a clear standard structure, the machine reads that as suspicious.
That is not a bug someone will fix next year. It is what the method does.
You do not need a critic to make this case. Turnitin makes most of it.
| Claim | Source | Figure |
|---|---|---|
| Document-level false positives, for documents with 20% or more AI writing | Turnitin | under 1% |
| Sentence-level false positives, meaning highlighted sentences that were human-written | Turnitin | about 4% |
| TOEFL essays by non-native speakers flagged as AI, average across seven detectors | Stanford, peer-reviewed | 61% |
| Those essays flagged by at least one of the seven detectors | Stanford, peer-reviewed | 98% |
| Best accuracy achieved by any of 14 detectors tested across 756 tests | Weber-Wulff et al., peer-reviewed | under 80% |
| Accuracy one detector advertised, against what regulators measured | US FTC order | 98% vs 53% |
Take the 4% figure and think about what it means on screen. When you open a flagged report, individual sentences are highlighted. By Turnitin's own arithmetic, roughly one in every twenty-five of those highlights is a sentence a human wrote. Nothing on the screen marks which ones. The wrong highlight looks exactly like the right ones.
You know which sentences you wrote. Your marker only knows what the report says. That asymmetry is the entire problem.
Turnitin also declines to show any score between 1% and 19%, and says why: false positives are more common in that range. So the company has drawn a line below which it does not trust its own output. Worth asking whether your institution knows that.
The Stanford paper is not the only one. A separate peer-reviewed study tested fourteen detection tools across 754 documents, and not one exceeded 80% accuracy. Only five got past 70%. The authors also found the tools were biased toward calling machine text human, which is the error nobody complains about.
This is the part worth quoting in an appeal, because it is institutions rather than critics making the argument.
Curtin University disabled Turnitin's AI writing detection across all campuses from 1 January 2026. The University of Waterloo switched its off in September 2025 after internal testing flagged human-written text as fully AI-generated. The University of Cape Town, Australian Catholic University, the Australian National University and Macquarie University have all done the same. Trackers now list more than fifty institutions.
Australian Catholic University logged around 6,000 integrity allegations in 2024, roughly 90% of its total caseload. About a quarter were dismissed.
Vanderbilt University disabled Turnitin's AI detector in August 2023 and did the multiplication out loud. At roughly 75,000 submissions a year, even a 1% false positive rate means about 750 students wrongly flagged. One percent is a small number until it has students attached to it.
In January 2026 a New York court ruled on this directly, and the details matter because they are so ordinary.
Orion Newby was a first-year student at Adelphi University, diagnosed with autism, enrolled in a university support programme for students with neurodevelopmental differences. He submitted an essay on Christianity and Islam for a World Civilizations class. Turnitin scored it as 100% AI-generated.
He had used Grammarly, and tutors the university itself provided. He said so when asked. He ran the essay through two other detectors, both of which called it human-written. The university upheld the violation, did not give him a copy of the report, and denied him an appeal. His family spent six figures on legal fees.
Justice Randy Sue Marber annulled the finding on 28 January 2026 and ordered the record expunged, describing the university's decision as without valid basis and devoid of reason.
Note what actually won it. Not a better detector score. The university's failure to follow a fair process, and its refusal to weigh evidence that contradicted the tool.
Newby is the case everyone cites, and citing it alone gives a misleading picture.
Haishan Yang, a doctoral student and non-native English speaker, lost. The Minnesota Court of Appeals upheld his expulsion in an opinion filed on 2 February 2026, and his separate federal claim had already been dismissed in October 2025. A Yale student was refused a preliminary injunction in May 2025, though his broader case continues. A case against the University of Michigan is still pending.
The pattern across all of them is consistent and it is worth understanding, because it tells you where to put your effort. Students win on process, not on the science. No court has yet ruled that detectors are invalid. What courts have ruled on is universities ignoring contrary evidence, denying a real appeal, or failing to accommodate a disability. So build your case around how you were treated, not around whether the tool is any good.
We built a writing checker, and it is important to say what it does not do.
It will not tell you whether your text was written by a human or a machine, because nothing can do that reliably. It will not give you a percentage to wave at your marker. It cannot clear your name.
What it does is show you which parts of your writing carry the patterns detectors react to. The words that cluster in machine text. Flat sentence rhythm. The hedging and the tidy uniform paragraphs. Each flag comes with an explanation of why it is flagged, and the rewriting is left to you, because a tool that rewrites your voice is a tool that makes the next accusation more likely.
It runs entirely in your browser. Your text is never uploaded, which matters when the text is the subject of a disciplinary process.
Used before you submit, it tells you where your writing reads as machine-made so you can decide whether to change it. That is a genuinely useful thing. It is not a defence, and we will not sell it as one.
Detectors measure predictability, not authorship. Plain writing, formulaic structure and English learned as a second language all read as predictable, which is why researchers turned a 5% flag rate into 57% by simplifying vocabulary alone. The companies publish their own error rates and say a score should not stand on its own. So do not fight the number. Show the work: version history, dated drafts, notes, and your own account of your argument.
See what reads as machine-made, free and private →Sources: Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. and Zou, J. (2023) "GPT detectors are biased against non-native English writers", Patterns 4(7), 100779, for the 61% average false-positive rate on TOEFL essays, the 98% flagged-by-at-least-one figure, and both vocabulary experiments. Turnitin's published guidance and its chief product officer's 2023 statements, for the document-level and sentence-level false positive rates, the withheld 1–19% band, and the position that a score is not a determination of misconduct. Vanderbilt University's August 2023 statement on disabling Turnitin's AI detector. Matter of Newby v Adelphi University, 2026 NY Slip Op 26021, Supreme Court of New York, Nassau County, decided 28 January 2026. Weber-Wulff, D. et al. (2023) "Testing of detection tools for AI-generated text", International Journal for Educational Integrity 19:26, for the fourteen-tool test and the sub-80% accuracy ceiling. Curtin University's statement on disabling Turnitin AI detection from 1 January 2026, and the University of Waterloo's September 2025 decision. Yang v University of Minnesota, Minnesota Court of Appeals opinion filed 2 February 2026, and Rignol v Yale University (D. Conn.), injunction denied May 2025. US Federal Trade Commission final order concerning Workado, August 2025. This article is general information about how these tools work and is not legal advice; if your case is serious, your student union or ombudsman is the right first call.