Skip to content
Sunday, August 23, 2026
Juvenile Pre & PostEducation & youth
Research · Learning · Evidence
Students

Why AI detectors still get student writing wrong

A Stanford study found detectors flagged 61% of non-native English writers' essays as AI-generated, and a growing list of universities is dropping the tools rather than trusting the score.

Why AI detectors still get student writing wrong

AI detectors still cannot reliably tell a student's own writing from a machine's, according to a Stanford study that found detectors falsely flagged 61% of essays by non-native English speakers as AI-generated, and a growing list of universities is dropping the tools rather than trusting the score.

How do AI detectors decide a paper was written by a machine?

Most detectors score a text's "perplexity" — a statistical measure of how predictable the word choices are — and flag writing that reads as unusually smooth or formulaic, according to a Stanford study led by professor James Zou. Predictable, low-perplexity phrasing is exactly what generative models produce, but it is also what careful, rule-following writers produce, especially students who learned English through formal instruction rather than immersion.

That overlap is the core design flaw. A detector built to catch statistically "too smooth" prose has no reliable way to separate a chatbot from a student who was taught to write in short, correct, unadorned sentences.

The Stanford team's numbers show how widely that flaw spread across the tools they tested: all seven detectors unanimously flagged 19% of the TOEFL essays as AI-written, and 97% of those essays were flagged by at least one of the seven. That is not one outlier product misbehaving — it is a pattern shared across the detector market the researchers sampled.

How often are AI detectors actually wrong?

It depends whose number you use, and that gap is the whole story. Turnitin says its detector runs "a less than 1% false positive rate" on its own test set, and its chief product officer has written that instructors should still "apply your professional judgment" because the rate is not zero. The Stanford researchers tested a different population — TOEFL exam essays from non-native English speakers — and found seven commercial detectors, including tools comparable to what schools use, misclassified 61.22% of those essays as AI-written, while scoring native-speaker eighth-grade essays with near-perfect accuracy.

Both numbers can be true at once. A detector's error rate on a vendor's general test set says little about its error rate on the specific students most likely to trigger it.

ClaimSourceWhat it measured
Under 1% false positive rateTurnitin, company blog, March 2023Turnitin's own internal test set, general population
61.22% false positive rateStanford study, May 2023Seven commercial detectors tested against TOEFL essays from non-native English speakers specifically

Neither figure is fabricated or in dispute; they simply answer different questions, and a school that only hears the first number is missing the second.

Which students are most likely to get falsely flagged?

English learners carry the largest documented risk, according to the Stanford findings and reporting in Education Week, which also found low-income students using school-issued devices face disproportionate scrutiny because their activity is more visible to monitoring software in the first place. That reporting adds a second layer to the problem: only 25% of teachers surveyed said they felt "very effective" at telling AI-written work from human work on their own, even as 68% had adopted detection tools — a gap between confidence and tool use that puts the burden of judgment on software teachers do not fully trust.

Stanford's Victor Lee put the stakes plainly in that same reporting: detectors "are fallible, you can work around them," and "an incorrect accusation is a very serious accusation to make."

Are schools still relying on these tools in 2026?

Fewer than before, and the retreat is accelerating. Inside Higher Ed reported this month that Yale, Vanderbilt, Johns Hopkins, and Indiana University now prohibit AI detection tools outright, citing false positives that have already prompted student lawsuits. The reporting describes an "unsustainable arms race": students adopt "humanizing" tools built to defeat detectors as fast as detectors improve, leaving instructors chasing a moving target rather than a fixed answer.

UC San Diego's Tricia Bertram Gallant, quoted in that reporting, argued the honest fix is not a better detector but a different kind of assignment: "the real trick is acknowledging that [20th-century assessments] don't work anymore." UCF's Kevin Yee called the moment "difficult" and "delicate," with no clean answer yet for instructors caught between suspicion and fairness.

Inside Higher Ed's reporting also notes that Turnitin's own detection software has produced unexpectedly high false-positive rates in practice — a detail worth sitting with given that Turnitin's own blog frames its rate as reassuringly low. The two claims are not necessarily contradictory, but the gap between a vendor's stated accuracy and an institution's observed accuracy is exactly what pushed Yale, Vanderbilt, Johns Hopkins, and Indiana to stop relying on the score at all.

What should a teacher do instead of trusting a single AI score?

Treat any detector output as one weak data point, never a verdict. Maine middle school principal Shawn Vincent, cited in the same Education Week reporting, described using detection scores alongside other evidence rather than as a standalone basis for discipline — a distinction that reporting's data suggests matters, since 40% of students in its survey faced discipline based partly on how they reacted when confronted, not on the score itself.

The schools moving away from detection entirely are betting on process instead of forensics. Inside Higher Ed's reporting lists the alternatives now spreading: assignments that require drafts and process documentation, in-person oral defenses of written work, proctored assessments for high-stakes writing, and smaller-group instruction where a teacher already knows a student's voice well enough to notice a real shift.

None of that requires abandoning AI tools in class — it requires not outsourcing the judgment call to a percentage on a dashboard.

Where this leaves classroom policy

For a school still running Turnitin or a similar detector: use the score as a prompt to look closer, not as proof, and weight it least for English learners and students already flagged by other monitoring software, per the populations Stanford and Education Week identify as highest-risk. For a school building new policy from scratch in 2026: the pattern in Inside Higher Ed's reporting — process documentation plus oral or in-person checks — is what peer institutions are adopting instead of a detector score. Neither Stanford's study nor Turnitin's own disclosure establishes that any detector is accurate enough to justify discipline on its own, and no source reviewed here claims otherwise.

Frequently asked questions

These answers are limited to what Stanford's research, Turnitin's own disclosures, and reporting in Education Week and Inside Higher Ed establish.

Is a low false-positive rate from a vendor enough to trust the tool?

Not on its own. Turnitin's under-1% figure comes from its own test set; Stanford's independent test on non-native English writers found a 61.22% false-positive rate for comparable detection technology, so the population being tested changes the answer.

Should a single flagged score lead to discipline?

The sources reviewed here say no. Turnitin's own chief product officer says instructors must apply judgment because the rate "is not zero," and Stanford's Victor Lee calls a wrongful accusation "a very serious" outcome to risk.

Are colleges banning these tools outright?

Some are. Inside Higher Ed reported in August 2026 that Yale, Vanderbilt, Johns Hopkins, and Indiana University now prohibit AI detection tools, citing false positives and resulting lawsuits.

What are schools doing instead?

Inside Higher Ed's reporting describes a shift toward process documentation, in-person oral defenses, proctored high-stakes writing, and smaller classes where instructors already know a student's writing voice.

Are English learners really flagged more often than other students?

Yes, on the evidence reviewed here. Stanford's test of TOEFL essays found a 61.22% false-positive rate for non-native English writers against near-perfect accuracy on native-speaker essays, and Education Week's reporting names English learners among the groups facing disproportionate scrutiny from detection software in practice.

For a related technology perspective, read AI detectors don't prove a student used AI to cheat.

Sources

  1. Stanford HAI — 'AI-Detectors Biased Against Non-Native English Writers'
  2. Turnitin — 'Understanding false positives within our AI writing detection capabilities'
  3. Education Week — 'More Teachers Are Using AI-Detection Tools. Here's Why That Might Be a Problem'
  4. Inside Higher Ed — 'AI Detectors Are Out, New Assessments Are In'