Skip to content
🎁
Need the full writing workflow?
Draft, translate, and refine English in one workspace.
Start for free
Authorship Certificate

How to Prove Your Essay Is Human-Written: A Guide for ESL Students

AI detectors flag non-native English writing at dramatically higher rates than native writing. If your university uses Turnitin or GPTZero, you need an authorship defense strategy before you submit — not after you get accused. Here is the step-by-step approach, grounded in peer-reviewed research and recent court rulings.
Alex Zhovnir
Alex Zhovnir
12 min read
May 2026
How to Prove Your Essay Is Human-Written: A Guide for ESL Students

In this article

🎁
Need the full Diglot workflow?
Keep drafting, translation, grammar review, and rewriting in one place.
Start for free

The short answer

You prevent a false-positive AI flag by building evidence of your writing process before you submit, not by arguing about your finished text afterwards. A peer-reviewed Stanford study published in Patterns (Cell Press) in 2023 found that seven commercial AI detectors incorrectly flagged 61.22% of human-written ESL essays as AI-generated, while classifying essays by native English writers with near-perfect accuracy. The finished essay cannot clear you. The record of how it was written can.

That record is also what settles disputes. In January 2026 the New York Supreme Court ruled in Matter of Newby v. Adelphi University that a university decision resting on a 100% Turnitin score was "without valid basis and devoid of reason," ordered the penalty rescinded and the record expunged. Below: the research behind the false-positive rate, the four court cases decided or pending so far, the universities that have switched AI detection off, and the eight-step evidence trail to set up before you write.

Why do AI detectors flag ESL writing?

Short answer: AI detectors do not detect AI. They measure how predictable and how uniform a text is, and careful non-native English is both.

AI detectors measure two statistical properties of text:

  • Perplexity — how predictable each word choice is. Low perplexity means the text follows common patterns.
  • Burstiness — how much sentence length and structure vary. Low burstiness means the sentences are uniform.

Large language models produce low-perplexity, low-burstiness text because they are designed to pick the most probable next word and maintain consistent output. When a detector sees text with those properties, it flags it.

Non-native English speakers also produce low-perplexity text. If you are still building vocabulary, you use safer word choices. If you studied grammar from textbooks, you write in standard patterns. If you translate from your native language, the result tends to be structurally uniform. To a detector, careful and correct English written by a second-language writer looks the same as machine output.

How often do AI detectors flag human ESL writing as AI?

Short answer: in a peer-reviewed Stanford test of seven commercial detectors, 61.22% of human-written ESL essays were flagged as AI-generated, and 97.8% were flagged by at least one of the seven.

In 2023, researchers at Stanford (Liang, Zou, et al.) published a study in Patterns (Cell Press) titled "GPT detectors are biased against non-native English writers". They ran 91 human-written ESL essays (from TOEFL preparation materials) and 88 essays written by US 8th graders through seven commercial AI detectors.

Measured on human-written essaysResult
ESL essays incorrectly flagged as AI-generated61.22%
ESL essays flagged by at least one of the seven detectors97.8%
ESL essays flagged unanimously by all seven detectors19.8%
Essays by native English writers (US 8th graders)Correctly identified with near-perfect accuracy

To prove causation, the Stanford researchers ran a follow-up test. They took verified-human native English essays and asked ChatGPT to "simplify" the language to sound non-native. The false-positive rate spiked. They took the ESL essays and asked ChatGPT to "enhance" them to sound native. The false-positive rate dropped. The detectors were not finding AI. They were punishing linguistic simplicity.

A 2026 study by Hadra, Cambridge, and Mesbah in the International Journal for Educational Integrity (Springer), "AI Detectors Fail Diverse Student Populations", found that Turnitin achieved only 61% overall accuracy, calling the trade-off between detection power and false accusations a "structural mathematical limit, not an engineering flaw that can be patched."

What have courts ruled about AI detector accusations?

Short answer: where a school's case rests on a detector score by itself, judges are increasingly sceptical; where independent evidence of cheating also exists, courts defer to educators.

Multiple lawsuits in 2025 and 2026 are testing whether schools can discipline students based on AI detector scores alone. Four cases show the pattern.

CaseCourtStatusWhat the accusation rested on
Doe et al v. Palo Alto Unified School District et al. (docket 5:25-cv-04202)U.S. District Court, N.D. Cal.PendingA 76% Turnitin score on an essay
Matter of Newby v. Adelphi UniversityNew York Supreme CourtStudent won, January 2026 — penalty rescinded, record expungedA 100% Turnitin score on a history paper
Doe v. YaleU.S. District Court, D. Conn.PendingA GPTZero flag on a final exam
The Hingham case (Massachusetts)Federal courtInjunction deniedAI-hallucinated citations, including a non-existent book — not the detector alone

Doe v. Palo Alto Unified School District — pending federal litigation

A family in California filed a federal lawsuit (Doe et al v. Palo Alto Unified School District et al., docket 5:25-cv-04202) after their child's essay was flagged at 76% by Turnitin. The student's grade dropped from a high B/low A to a C. The family submitted a 1,162-page evidentiary packet — revision history, drafts, Google Docs timestamps — showing the work was written by hand over days. The district declined to restore the grade. The case is now pending in the U.S. District Court for the Northern District of California, with claims filed under Title VI (national origin discrimination), Title IX, and Due Process.

Newby v. Adelphi — the student won

In January 2026, the New York Supreme Court ruled in Matter of Newby v. Adelphi University. Orion Newby, a freshman with Level 2 Autism Spectrum Disorder enrolled in Adelphi's "Bridges Program," was flagged at 100% by Turnitin on a history paper. The university upheld a failing grade. His family spent over $100,000 in legal fees.

The judge in Newby v. Adelphi ruled the university's decision was "without valid basis and devoid of reason" and ordered the penalty rescinded and the record fully expunged. That ruling establishes — at least at the state-court level in New York — that an AI detector score alone is not a valid basis for academic punishment.

Doe v. Yale — ESL and Title VI, pending

In Doe v. Yale, an Executive MBA student at Yale was suspended for a year after a final exam was flagged by GPTZero. The student, a non-native English speaker, filed suit in U.S. District Court for the District of Connecticut arguing that the detector's algorithm has implicit bias against ESL writers — the exact mechanism the 2023 Stanford study documented — and that using it to discipline a non-native speaker amounts to national-origin discrimination under Title VI. The case is pending.

The counter-example: the Hingham case

Not every case goes the student's way. In the Hingham case in Massachusetts, a federal court denied a student's injunction after the school found AI-hallucinated citations in the student's work — including a non-existent book. The court reasoned that the school's case rested on independent evidence of cheating, not the detector alone.

For an ESL student doing their own work, the practical implication of those four cases is straightforward: the detector score is not what decides the outcome. Preserved evidence of your writing process is.

Which universities have turned AI detection off?

Short answer: Waterloo, Vanderbilt, MIT and Curtin have publicly disabled or restricted Turnitin's AI detector, citing unreliability and bias against non-native English speakers, and a longer list of institutions has quietly switched the feature off inside their learning management systems.

InstitutionWhat it didStated reason
University of WaterlooFormally discontinued Turnitin's AI detection in September 2025The tools are "unreliable" and "biased toward students whose first language is not English"
VanderbiltDisabled the detectorFalse positive risks and ESL bias
MITPublished internal teaching guidance titled "AI Detectors Don't Work. Here's What to Do Instead."
Curtin University (Australia)Announced it will disable Turnitin AI detection starting in 2026
American University, Boston University, UC Berkeley, Colorado State, DePaul, Georgetown, Michigan State, NYU, University of Cape TownQuietly disabled or restricted AI detection through their learning management systemsMost did this without press releases

Even Turnitin itself has shifted. In February 2026, Turnitin CEO Chris Caren announced a pivot "from detection to transparency." The company's own documentation acknowledges an uncertainty band of roughly plus or minus 15 percentage points and states the AI indicator "should not be the sole basis for punitive action."

How do you prevent a false-positive AI flag before you submit?

Short answer: write in a tool that keeps timestamped revision history, keep your outlines, native-language drafts and research notes, and spread the work across several sessions — so the evidence exists before anyone asks for it.

Court rulings and university policy reversals will not help you in next Tuesday's meeting. What matters in practice is the evidence you already have when someone questions your writing. The eight steps below cover what to set up before you write, what to track while writing, and how to respond calmly with evidence if a detector flag arrives after submission.

Before you write

1. Choose a writing tool that keeps history. Google Docs preserves a full version history by default — every edit, paste, deletion, and timestamp. So does Microsoft Word (with AutoSave on OneDrive), Notion, and most modern editors. If your tool does not keep revision history, switch to one that does. This is the foundation of any authorship defense.

2. Keep your native-language drafts. If you outline or draft in your first language before translating into English, keep that original. A paper trail showing "here is my outline in Mandarin, here is the English translation, here are four rounds of revision" is evidence that no detector can contradict.

3. Save research notes separately. Keep a running document (or folder) of your sources, quotes, notes, and reasoning. If someone asks how you arrived at a particular argument, you can show the path from source to draft.

While you write

4. Write in sessions, not in one sitting. Revision history that shows writing spread over multiple days — with breaks, revisions, deleted paragraphs, restructured sections — is the strongest evidence of human authorship. A document that appears fully formed in one session (even if you genuinely wrote it that way) is harder to defend.

5. Do not paste large blocks from external sources. If you draft sections in a separate document and paste them into your final paper, the version history shows a sudden appearance of fully-formed text. Draft directly in the submission document when possible, or at minimum keep both documents with their own version histories.

After submission — if you are flagged

6. Do not apologize or admit fault. The instinct when a professor asks "did you use AI?" is to over-explain. Lead with evidence instead. Calmly present your revision history, native-language drafts, research notes, and timestamps. Reference the 2023 Stanford study if you are an ESL writer — it is peer-reviewed and directly on point.

7. Ask specific procedural questions. Request in writing: which detector was used, what score triggered the flag, what your institution's formal procedure is for handling AI accusations, and what your appeal options are. The strongest cases against schools involve procedural due process violations — institutions that did not give students fair notice, a chance to present counter-evidence, or a clear appeal path.

8. Know your institution's policy. Read your school's academic integrity policy before any dispute arises. If it does not address AI specifically — or does not specify what evidence counts as a defense — that is an argument worth raising. Many institutions adopted detector tools faster than they updated their integrity codes.

What are teachers' unions and legislatures doing about false AI flags?

Short answer: the NEA has stated that biased AI cheating detection has incorrectly flagged students, including multilingual learners; the AFT has passed a resolution urging safeguards for individual rights in AI deployment; and states including Idaho and California are legislating limits on automated decisions in education.

The NEA (National Education Association) — representing 3 million educators — published Five Principles for AI in Education stating that "biased AI cheating detection applications have incorrectly flagged students for misconduct" and specifically naming "emergent multilingual learners" who "have been falsely accused." The American Federation of Teachers passed a resolution urging safeguards for individual rights in AI deployment.

At the state level, Idaho enacted SB 1227 — a statewide AI-in-education framework requiring human-centered oversight and prohibiting high-stakes automated decisions. California's AB 1159, which passed its first chamber, would prohibit ed-tech vendors from training models on student submissions. Maryland and Illinois have similar bills pending.

The direction is clear: the era of treating a detector score as a verdict is ending. But the transition is slow, and individual students bear the risk in the meantime.

How does Diglot prove you wrote it?

Short answer: Diglot's Authorship Certificate signs every edit, paste and AI-assist in your document and produces a public verification URL that anyone can open — no Diglot account required to verify.

Diglot is a bilingual writing workspace built for the exact workflow that gets ESL students flagged: thinking in one language, writing in another, and refining until the English is clean.

The Authorship Certificate tracks every edit, paste, and AI-assist on your document and signs the resulting event chain with a cryptographic key. The result is a public verification URL that anyone — your professor, your department, an appeals committee — can open in any browser to see a tamper-proof record of how the document was actually written.

If you are an ESL student who drafts in your native language and publishes in English, the Certificate captures that entire journey: the first draft, the translation, every round of revision. When someone questions whether you wrote your own work, you do not need to argue. You send a link.

Read the companion articles for more context: AI detection lawsuits in 2026, whether Turnitin's AI detection is accurate, how common AI detector false positives are, why AI detectors misread non-native English, how to make your English writing sound natural, and what to do if a client flags your writing.

Start a free draft with Diglot

Frequently asked questions

The questions below cover what ESL students ask most often about false AI flags: how to prevent one before submission, whether expulsion is realistic, how accurate detectors actually are on non-native English writing, whether grammar and paraphrasing tools increase risk, what evidence holds up in an appeal, and what to do specifically if your university still relies on Turnitin's AI detector to flag essays.

How do I prevent a false-positive AI flag before I submit?

Build the evidence trail while you write, not after a flag. Draft in a tool that keeps timestamped revision history, keep the outlines and native-language drafts you started from, save research notes and source links separately, write across several sessions rather than one sitting, and avoid pasting large blocks of finished text in from other documents. Preserved process evidence is what holds up when a detector score does not.

Can I get expelled for a false AI detection?

Penalties vary by institution — from a warning to academic misconduct charges. However, courts are increasingly skeptical of punishments based solely on detector scores. In Newby v. Adelphi, a judge overturned the penalty entirely. Your best protection is preserved evidence of your writing process.

Are AI detectors accurate for ESL writing?

No. A peer-reviewed Stanford study found that seven commercial AI detectors incorrectly flagged 61.22% of human-written ESL essays as AI-generated. The bias is structural: detectors measure text predictability, and non-native English writing is naturally more predictable due to constrained vocabulary and standard grammar patterns.

Does using Grammarly or a paraphrasing tool make my essay look like AI?

Grammar and paraphrasing tools can reduce perplexity (word predictability) and burstiness (sentence variation) — the same signals detectors flag. Using them does not mean you cheated, but it can increase your false-positive risk. Keep records of your drafts before and after tool use.

What evidence do I need to prove I wrote my essay?

The strongest evidence includes revision history with timestamps (Google Docs, Word, Notion), native-language drafts or outlines, research notes and source links, and intermediate draft versions showing the work evolving over multiple sessions. In the Palo Alto federal case the family submitted a 1,162-page packet of revision history, drafts and Google Docs timestamps. Evidence captured before submission is stronger than anything assembled after a flag.

My university still uses Turnitin. What should I do?

Write in a tool that preserves revision history, keep your research notes, and if you draft in your native language, save those drafts. Build the proof trail before you submit — not after a flag. It also helps to know that Waterloo, Vanderbilt, MIT and Curtin have publicly disabled or restricted Turnitin AI detection, and that Turnitin's own documentation states the AI indicator "should not be the sole basis for punitive action." Read the full legal landscape article for context on how universities are responding.

Sources

This guide draws on peer-reviewed research from Stanford and Springer, federal and state court rulings, public disabling statements from major universities, position papers from the NEA and AFT, and Turnitin's own disclosures. Each citation is linked below. Where rulings are still active or institutions have not yet finalised policy, we will revise this page as material changes land.

This article references active litigation. Case statuses may change after the update date shown at the top of this page.

Write in your language,
publish in English

Move from rough bilingual drafts to clearer English in one connected writing workflow.

Start for free

*No credit card required

Diglot.ai - bilingual writing tool, write and translate in one app

Other Blogs

We carefully select blog topics so that you can get the most useful and precise information