In this article
Three different answers, and the confusion in every forum thread comes from mixing them up.
Technically, no AI detector knows which tool touched your text. Detectors read the text. Turnitin says so directly in its own AI-detection FAQ: its detector is not tuned to target Grammarly’s spelling, grammar and punctuation changes.
By policy, at some universities Grammarly absolutely counts — it is named in the rules, with a disclosure template you are expected to fill in. At others it is not. The policies disagree with each other, in writing.
In practice, students have been penalised after a Turnitin flag while saying they used nothing but basic corrections. The best-known case cost a student her scholarship.
All three are true at once. Below is what each one actually says, with the source and the date, and what to do about the situation you are actually in — which is not “how do I pass the scan” but “how do I show this is mine”.
The question is not the tool. It is the function.
This is the single most useful reframe available, and it is the rare point where the detector vendor and the editor vendor say the same thing.
| What you did | Flagged? | Who says so |
|---|---|---|
| Accepted spelling, grammar, punctuation fixes | Generally not | Turnitin FAQ; Grammarly docs |
| Accepted a rewritten sentence or paragraph | Expected to raise the score | Turnitin FAQ; Grammarly docs |
| Used draft generation, paraphrasing, summarising | Likely flagged | Turnitin FAQ |
| Ran the text through a paraphraser or spinner | Detected as likely paraphrased | Turnitin FAQ |
Turnitin’s FAQ (checked live on 23 August 2026) states that its detector is not tuned to target Grammarly-generated spelling, grammar and punctuation modifications, and that in tests on human-written documents, changes made by Grammarly free and premium were mostly not flagged. The same answer then draws the line explicitly: content produced by Grammarly’s generative features — draft generation, paraphrasing, summarising — will likely be flagged.
Grammarly’s own support documentation agrees. It states that traditional, non-generative corrections in the form of its red and blue underlines should typically not affect the percentage score, and that using its generative assistant to meaningfully rewrite sentences and full paragraphs will raise it.
So: “does Grammarly count as AI” has no single answer because Grammarly is not one thing. Accepting a comma and accepting a rewritten paragraph are different acts, and everyone involved — including the company selling you the tool — agrees they are.
The numbers behind that line
Copyleaks ran the comparison in May 2024. A thousand files were corrected for correctness and clarity, excluding the generative platform; all 1,000 came back as human content. Five hundred of the same original files were then run through Grammarly’s generative “Improve” feature; 31.6% came back as AI. Note what this is: a detector vendor testing its own detector. It is directional evidence, not an independent result.
Originality.AI ran a smaller version in 2025 on ten human IELTS essays, splitting them into light edits (grammar suggestions only) and heavy edits (accepting rephrasing and rewriting). Light edits mostly stayed human; heavy edits significantly changed whether the content was flagged. Same shape, same caveat, and n=10.
There is, as far as we can find, no independent, non-vendor, reproducible experiment on rule-based grammar checking versus AI detectors. The academic work that exists — including the APT-Eval study of AI-polished writing — tests edits made by a language model, not by a grammar checker. That distinction matters, and nobody selling you a detector will draw it for you.
Where the fear comes from: the case people actually remember
Marley Stevens, a student at the University of North Georgia, submitted a paper on 3 October 2023 and received a zero on 26 October after a Turnitin AI flag. She lost her scholarship. In February 2024, after a hearing, she was placed on conduct probation for a year and required to pay for a seminar.
Her account is that she used the free version of Grammarly for punctuation and spelling only. The university has not published its version, and that asymmetry is worth stating plainly rather than glossing over: what is documented publicly is the student’s account, and it received wide coverage.
What makes the case land is not the verdict. It is that the process offered no way to demonstrate the opposite. A finished document is thin evidence in either direction. The student had nothing else to show, and neither would you — which is the same structural problem behind every case in our roundup of AI-detection lawsuits.
What detectors actually punish — and why “write simpler” is bad advice
The mechanism is real and it is measurable, but it points the opposite way from what most advice assumes.
Liang and colleagues tested seven commercial detectors and published in Patterns in July 2023. Essays by US eighth-graders were classified almost perfectly. TOEFL essays were not: the detectors misclassified over half of them, at an average false-positive rate of 61.22%. The essays that all seven detectors got wrong shared one measurable property — significantly lower perplexity, meaning more predictable word choice.
Then they ran the interventions, and this is the part that gets left out of every summary:
- Enrich the vocabulary. Prompting the TOEFL essays toward more native-like word choice dropped the average false-positive rate from 61.22% to 11.77%. Perplexity went up; flags went down.
- Simplify the vocabulary. Running the reverse prompt on the US eighth-graders’ essays pushed misclassification from 5.19% to 56.65%. Perplexity went down; flags went up.
Read those two lines together and the popular theory collapses. Detectors do not punish “text that has been cleaned up”. They punish predictable, templated, impoverished language. A grammar checker that fixes your commas does not make your word choice more predictable. A tool that replaces your specific verb with the most statistically expected one does.
Which also means the advice you sometimes see — write a bit worse so you sound human — is precisely backwards. It makes you more likely to be flagged, not less.
One honest gap: we could not find any study measuring what happens to perplexity after a rule-based grammar checker specifically. The chain “grammar checker → lower perplexity → more flags” is a plausible hypothesis, not a demonstrated fact, and we are not going to state it as one.
The double bind nobody warns multilingual writers about
Here is the trap, written into an actual university policy.
The University of York’s guidance on AI and translation tools (published November 2023) defines false authorship as work created or modified using undisclosed help to the point where you can no longer be considered the author. Among the signs it lists, under «enhancement beyond expectations»: work that is consistently grammatically flawless and highly cohesive may suggest false authorship — particularly, it adds, when similar work cannot be produced under conditions where you do not have those tools available. York states plainly that it expects your work to reflect what it already knows of you from class participation, homework and day-to-day interaction.
Sit with that for a second if you write English as a second language. If the feeling has a name, it is flagxiety — and this is the mechanism underneath it.
Leave the draft as it comes out, and you risk the low-perplexity problem Liang measured. Polish it until it is clean, and you risk being flagged for prose that is too good for you — a judgement made against your own previously demonstrated level. Both doors are alarmed. That is not a fear you invented; it is written down in the policy.
There is exactly one thing that resolves the contradiction, and it is not a better editing strategy. It is evidence of how the document was made.
Google Translate and DeepL: the policies genuinely disagree
Do not let anyone tell you there is a universal rule here. There is not, and the disagreement is the story.
Named as an AI tool requiring disclosure:
- Sheffield Hallam University publishes an AI Transparency Scale that puts DeepL and Grammarly on the same ladder. Its own examples include using DeepL to shape an initial draft by translating ideas from Spanish, and using DeepL and Grammarly together to translate and refine an assignment. Its translation guidance says you should not submit AI-translated work without acknowledging it.
- McMaster University names Google Translate, DeepL and Microsoft Translator in its guidance on translation and editing tools. For editing tools it lists four uses — grammar and spelling, readability, rewording or summarising, tone and style — and states directly that uses 2 through 4 could be considered academic dishonesty. It also supplies a disclosure template naming the tool and version.
Treated as a separate category:
- University of the Arts London states in its staff guide that although machine translation tools are driven by AI, they are different from generative AI tools that create content.
- University of York lists translation tools — Google Translate, Youdao, Baidu — as their own category, distinct from generative AI.
And Turnitin has no position at all: the word “translat” does not appear anywhere in its AI-writing-detection FAQ. Its detector processes long-form English, Spanish and Japanese; that is the whole statement. Anyone claiming Turnitin does or does not treat translators as AI is making it up.
Does translating raise your flag risk?
One study says yes, substantially. Originality.AI, June 2026: 498 human and 498 GPT-4o samples, translated through Google Translate out of English and back. Human English that started with a 0.40% false-positive rate came back at 28.02% after the round trip.
Two caveats belong permanently attached to that number. First, it is a round trip — English out, English back — and not the direction most multilingual writers actually work in, which is first language to English. Direct single-pass translation scored far lower in the same study. Second, it is a detector vendor testing its own detector. On what a human reader actually notices in translated prose — a different question from what a detector notices — we have a separate guide.
We are including the caveats because leaving them off would make this page like every other page on the topic. The number is still worth knowing.
What the current numbers actually are
The 61.3% figure travels everywhere without its date attached, and that is a problem for anyone relying on it.
It is from 2023, measured on 2023-era detectors, on 91 TOEFL essays sourced from a forum, compared against essays by American eighth-graders. Two vendors have contested it:
- Originality.AI argues the comparison group is a confounding variable and reports 5.04% on 1,607 IELTS essays.
- Pangram reports 0% on the same 91-essay TOEFL set and 0.012% across 25,021 ESL samples drawn from four corpora.
Both are vendors marking their own homework, exactly like Turnitin’s claim of under 1%, GPTZero’s 1.1% on TOEFL essays, and Grammarly’s own detector caveats.
The most useful recent number is independent. Researchers at Notre Dame published a controlled study in August 2026 on English-language abstracts across four disciplines. Their findings:
- Unmodified recent abstracts were flagged at 9–15%.
- Light editing that complied with the guidelines was flagged at 38–80%.
- Non-STEM writing fared significantly worse than STEM.
- Text run through a commercial “humanizer” left fewer than 4% of AI-labelled rewrites still flagged.
Their conclusion is stated flatly: detector scores should not serve as standalone misconduct evidence.
Look at those last two bullets together. Honest, permitted editing gets flagged between 38 and 80% of the time. Deliberate laundering gets through more than 96% of the time. The system as it currently works penalises the honest and rewards the evasive. That is not a reason to become evasive. It is a reason to stop treating a detector score as the thing that decides.
It is also why at least one major institution stepped back: Vanderbilt disabled Turnitin’s AI detector in August 2023, citing the bias against non-native English speakers and the fact that at 75,000 papers a year, a claimed 1% false-positive rate still means around 750 papers wrongly labelled. As of 2026 it still offers no institutionally supported AI detection tool.
So what do you actually do?
Not “beat the detector”. Every path in that direction ends at a humanizer, and the Notre Dame numbers show exactly what that market is: a tool that works better for people cheating than for people who are not. We are not going to point you there and we are not going to hint at it.
The move that actually resolves the double bind is different in kind. Stop offering the finished text as your only evidence, because it is bad evidence. Every detector reads the same finished text and reaches a different conclusion — that is what all the numbers above are measuring. A record of how the document came into existence is a different category of evidence entirely: it can be checked.
Two products exist in that category.
Grammarly Authorship got there first and deserves the credit. Grammarly has also said publicly that no AI detector can conclusively determine whether AI was used, and that detectors have been found to be biased against non-native English writers — a notable admission from a company that sells one. It positions itself against detectors precisely on this point: detectors analyse the final text, Authorship records the writing process. It works in Google Docs, Microsoft Word, the Grammarly Editor and Canvas, and it labels text as typed, pasted, AI-generated, AI-edited or grammar-corrected. It has one documented weakness: in Google Docs it depends on the browser extension and cannot see what happens outside it, so text produced in a desktop editor and pasted in shows as unknown — one review produced a document reading as 75% typed by a human that was not.
Diglot records the writing process as a signed, append-only event chain and issues an Authorship Certificate that anyone can verify without an account. Same category, and we built it because this exact problem is the reason the product exists — it is a bilingual writing tool for people who are structurally more likely to be falsely accused. We have an obvious interest in you finding this convincing, so treat this paragraph accordingly and go look at how verification works yourself.
Neither one proves innocence on its own, and any vendor telling you otherwise — us included — is overselling. What a process record does is change the conversation from a number you cannot interrogate to a history someone can actually look at.
Meanwhile, three things worth doing regardless of which tool you use:
- Read your own institution’s policy, and read it for the function rather than the brand. McMaster’s four-use breakdown is a better mental model than any list of banned apps.
- Keep your drafts. Version history, dated files, whatever your workflow already produces. It is not a certificate, but it is not nothing.
- Do not deliberately write worse. The Liang numbers are unambiguous on this, and it is the most common piece of bad advice in circulation. If you want to know which specific features of your own draft read as machine-like, our free why-flagged-as-ai checker names them without giving you a percentage — because a percentage is exactly the thing this article argues you should stop trusting.
Fact-checked 23 August 2026. Every claim on this page is sourced to a named document with a date: Turnitin’s AI-detection FAQ, Grammarly’s AI Detector user guide and 2024 authorship announcement, Copyleaks (May 2024), Originality.AI (2025 and June 2026), Liang et al. in Patterns (July 2023), the Notre Dame study (August 2026), and the published policies of Sheffield Hallam, McMaster, University of the Arts London, York and Vanderbilt. The Turnitin FAQ, the York policy and the Grammarly documentation were opened and read directly on 23 August 2026 rather than taken from a cache. Turnitin revises that FAQ often and shows only a relative update date, and much of the Grammarly section rests on its current wording — so this page is on a quarterly review cycle.
