In early August 2026, YouTuber Hank Green paused his channels after viewers noticed the phrase “I appreciate the pushback” in a Complexly video and concluded a chatbot had written the script. Green said he had used ChatGPT for research but that the line was his own. The episode is a public version of something that now happens quietly every day: people being falsely accused of using AI on work they wrote themselves.
Students, freelancers, marketers and employees are all facing it. And the uncomfortable answer is that almost none of the “evidence” people rely on actually proves anything. Here is what does, what does not, and what to do if it happens to you.
Why everyone suddenly thinks they can spot AI writing
The suspicion is not baseless. Writing really has changed since ChatGPT launched.
A 2025 study published in Science Advances by Dmitry Kobak and colleagues analysed more than 15 million biomedical abstracts indexed in PubMed between 2010 and 2024. The researchers borrowed the “excess mortality” method from epidemiology and applied it to vocabulary, measuring which words appeared far more often after 2022 than the pre-ChatGPT trend predicted. Words such as delve, intricate, meticulously, realm, pivotal and showcasing spiked sharply. The authors estimated that at least 13.5% of 2024 abstracts were processed with a large language model, rising to around 40% in some subgroups of journals and countries.
That is a real, measurable fingerprint. But it is a fingerprint on a population, not on a person.
Key takeaway: a statistical shift across millions of documents tells you nothing reliable about the single paragraph in front of you.
This is the trap almost every accusation falls into. A reader learns that AI overuses a word, spots that word once, and treats a population-level pattern as individual proof. Humans used “delve” and em dashes long before 2022, and plenty still do.
What actually gives away AI writing?
Almost nothing does on its own. The strongest genuine signals are contextual rather than stylistic: an unedited prompt or chatbot reply left inside the text, citations and statistics that do not exist when you check them, or an abrupt change from a writer’s established voice. Word choice, em dashes and tidy structure are hints at best.
It helps to separate the two tiers of evidence.
Weak signals (suggestive, never conclusive):
- Frequent em dashes, or words like delve, robust, tapestry, leverage and navigate
- Stock phrases such as “it’s important to note” or “let’s dive in”
- A rigid structure: short intro, three to five neatly balanced points, tidy conclusion
- Flawless grammar with no personal detail or specific example
- A score from an AI detector
Strong signals (worth investigating):
- Leftover conversational artefacts, such as “Certainly! Here is a revised version” or a model refusing a request mid-paragraph
- Fabricated sources: quotes, studies or page numbers that do not exist when verified
- Confident claims about events after the tool’s knowledge cutoff, stated wrongly
- No draft history at all for a long piece of writing
- The writer cannot explain their own argument when asked
Notice that the strong signals are all checkable facts. The weak signals are all taste.
Do AI detectors actually work?
Not reliably enough to accuse anyone. OpenAI withdrew its own AI Text Classifier in July 2023, citing a low rate of accuracy. In the company’s own evaluation, the tool correctly identified just 26% of AI-written text as “likely AI-written” while wrongly labelling 9% of human writing as AI-generated. That is roughly one false accusation in every eleven honest documents.
Commercial detectors that remain on the market have not solved the underlying problem, because there is no watermark to find. They are estimating how “predictable” text looks, and predictable prose is something careful human writers produce too.
Institutions have started to act on that. Vanderbilt University disabled Turnitin’s AI detector on 16 August 2023, and its teaching centre laid out the arithmetic plainly: Vanderbilt submitted around 75,000 papers to Turnitin in 2022, so even the vendor’s claimed 1% false positive rate would mean roughly 750 student papers incorrectly flagged in a single year. Several other universities have since restricted or switched off AI detection for the same reason.
Key takeaway: a detector score is a probability estimate, not evidence, and no serious institution should treat it as the sole basis for a penalty.
Why non-native English speakers get flagged most
This is the part of the story that deserves far more attention than it gets.
A Stanford team led by Weixin Liang published “GPT detectors are biased against non-native English writers” in the journal Patterns in 2023. They ran seven commercial detectors over two sets of genuinely human writing: essays by US eighth-grade students, and TOEFL essays by non-native English speakers. The detectors were close to perfect on the American school essays. They misclassified more than half of the TOEFL essays as AI-generated.
The likely cause is technical, not political. Detectors lean on perplexity, a measure of how surprising each word is given the previous ones. Writers working in a second language often use a more constrained, more conventional vocabulary, which produces low perplexity text that looks statistically like a language model’s output.
The practical effect is that an international student in Melbourne, a graduate applicant in Delhi, a job seeker in Warsaw and a contractor in Manila all carry a higher risk of being falsely accused of using AI than a native speaker writing the same quality of work. The researchers also showed the bias can be dodged simply by prompting a model to write more elaborately, which means detectors punish the honest writers and wave through the ones actively trying to cheat.
What to do if you are falsely accused of using AI
Stay calm and make the conversation about evidence rather than vibes.
- Ask exactly what the evidence is. A detector percentage, a flagged phrase and “it reads like AI” are three very different claims. Get it in writing.
- Produce your version history. Google Docs version history, Microsoft Word’s document history, Notion page history or Git commits show a document growing over hours. This is the single most persuasive defence available.
- Offer to explain the work out loud. Walk through your argument, your sources and why you cut a section. Anyone who genuinely wrote something can do this; it is far harder to fake than prose style.
- Cite the research. The Stanford Patterns study, OpenAI’s withdrawal of its own classifier and Vanderbilt’s decision are all public and easy to link.
- Do not rewrite in a panic. Rushed edits look like tampering. Keep the original file untouched.
- Follow the formal appeal route. Most universities and employers have one, and written appeals travel better than corridor conversations.
How to protect yourself before it happens
Prevention is mostly about leaving a trail.
- Draft in a tool that keeps automatic version history rather than typing straight into a submission box
- Keep your research notes, bookmarks and highlighted PDFs
- Read your institution’s or employer’s actual AI policy, and disclose assistance where it asks you to
- Keep concrete detail in your writing: a named example, a number you looked up, an anecdote only you would have
- If you do use AI for research, verify every fact independently before it reaches the page
Disclosure norms are still forming, and they differ sharply. Using a chatbot to summarise ten papers is acceptable at many workplaces and prohibited in many classrooms. Knowing which room you are in is most of the battle. We covered a related myth in our piece on whether AI really rejects your CV before a human sees it, and the pattern is the same: the technology is real, the folklore around it is mostly wrong.
The bigger picture
AI systems are genuinely getting more capable. Only last week the debate was about whether a model had solved a decades-old open maths problem. But capability and detectability are different questions, and detection is losing badly. As models get better at writing like people, the statistical gap that detectors depend on keeps shrinking.
That leaves process, not forensics, as the workable answer. Version history, oral defence, verified sources and honest disclosure all survive the next model release. A list of suspicious words does not.
Frequently Asked Questions
Can a teacher fail you based only on an AI detector score?
Most institutions now say no. Detector scores are probability estimates with documented false positive rates, and universities including Vanderbilt have disabled the tools for that reason. Policies vary, so check your own institution’s academic integrity rules, which usually require corroborating evidence before a penalty.
Does using ChatGPT for research count as cheating?
It depends entirely on the policy you are working under. Many workplaces treat it like using a search engine, while many courses require disclosure or ban it for assessed work. Read the specific rule, and when in doubt, disclose what you used and how.
Do em dashes mean text was written by AI?
No. Em dashes are a normal punctuation mark used by human writers for centuries, and their prevalence in AI output varies by model and version. Punctuation alone is one of the weakest possible signals and should never be treated as proof.
Are AI detectors biased against non-native English speakers?
Yes, and it is documented. A 2023 Stanford study in Patterns found seven commercial detectors misclassified more than half of TOEFL essays written by non-native English speakers as AI-generated, while being near-perfect on native-speaker school essays.
What is the strongest proof that something was AI-generated?
Contextual evidence, not style. Leftover chatbot text, fabricated citations that fail verification, and a complete absence of drafting history are far more telling than word choice. An author who cannot explain their own reasoning is also a stronger signal than any detector score.

