
AI Summary
AI detection tries to decide whether text was written by a human or a language model, usually as a probability score from a classifier. No publicly available detector is reliable enough to act on: it flags careful human writers as machines and passes lightly edited AI text, which is why acting on the number alone is indefensible.
- Detectors score statistical texture, not authorship, so predictable human prose reads as machine written.
- A short paraphrase pass drops the score under the threshold, making evasion trivial for anyone motivated.
- OpenAI retired its own classifier within months, citing low accuracy.
- Non native English writers are flagged at disproportionately high rates, which makes acting on scores unfair as well as unreliable.

AI detection is the attempt to determine whether a piece of text was written by a human or a language model, usually via a classifier that outputs a probability score. Here is the part vendors won't lead with: no publicly available detector is reliable enough to act on, and the confident percentage scores they display have already cost real people grades, jobs, and clients.
The reliability problem, stated bluntly
These tools measure statistical texture, how predictable the word choices are, how uniform the sentence rhythm is. That correlates with machine authorship; it does not establish it. Careful human writers, technical writers, and non-native English speakers all produce "predictable" text and get flagged for it. Meanwhile a light paraphrase pass takes machine text right back under the threshold. OpenAI, the company with the best possible visibility into its own models' output, launched a detection classifier in 2023 and shut it down within months, citing its low rate of accuracy. That should calibrate your trust in third parties claiming to solve, at 99%, a problem the model's own maker walked away from.
The false-positive consequences are not hypothetical. Students have been hauled into misconduct hearings on the strength of a score. Freelance writers have lost contracts because a client pasted their work into a free checker. There is no appeal process against a number, and that is exactly why acting on the number alone is indefensible.
Detector claims vs. documented reality
| The claim | The documented reality |
|---|---|
| "99% accurate" | Vendor-measured on their own curated test sets. Independent testing on edited, mixed, or out-of-domain text shows substantially worse performance, and accuracy claims quietly exclude the paraphrase case |
| "Near-zero false positives" | OpenAI retired its own classifier for low accuracy; Turnitin publicly added caveats about false positives after launch claims. Every large-scale deployment has produced wrongly flagged humans |
| "Robust to editing and paraphrasing" | Peer-reviewed work has repeatedly shown paraphrasing, by tool or by hand, collapses detection performance. Evasion is a solved problem for anyone motivated |
| "Detects all AI models" | Classifiers are trained on the output of yesterday's models and degrade against newer ones. The target moves every release; the detector retrains after |
| "Fair across all writers" | Researchers found detectors flag non-native English writers at disproportionately high rates, simpler vocabulary and regular structure read as "machine" to a perplexity-based classifier |
| "Google uses AI detection to penalize sites" | No evidence, and it contradicts Google's stated position that quality, not production method, is what's evaluated. See AI content for what the policies actually enforce |
How to check a detector before you trust it
- Feed it guaranteed-human text. Your own pre-2020 writing, old published articles, a chapter of public-domain prose. Every flag it raises here is a measured false positive rate you collected yourself.
- Feed it lightly edited AI text. Generate a passage, spend three minutes rewording it, rescan. Watch the score fall through the floor, that is the evasion cost for anyone who cares to evade.
- Run the same text through three detectors. When they disagree wildly, and they will, you have learned what a "97% AI" score is worth.
- Test your actual content domain. Detectors behave differently on technical docs, product copy, and essays. A tool benchmarked on student essays tells you nothing about your changelog.
- Decide in advance what a score would change. If no score would justify firing a writer or rejecting a submission on its own, and it shouldn't, be honest about what you're buying the tool for. The detection study on whether Google can identify AI writing is a useful sanity check on how thin the signal really is.
Common mistakes and fixes
- Firing or failing someone over a score. The tool cannot carry that weight. Fix: treat a flag as a reason to have a conversation and look at drafts, revision history, and subject knowledge, evidence a human can evaluate.
- Using "0% AI" as a quality gate for publishing. A page can score fully human and still be worthless; a genuinely useful page can flag. Fix: review content against reader value, the SEO implications of detection piece covers why provenance is the wrong axis.
- Paying for "undetectable AI" rewriting services. You are paying to make text statistically weirder, not better, and often making it worse to read. Fix: spend the same budget on editing and original input.
- Publishing detector scores as proof of anything. "This article is 100% human (verified by X)" impresses nobody who understands the tools and misleads everyone else. Fix: demonstrate humanity with bylines, experience, and specifics instead.
- Assuming detection will mature into reliability. The economics run the other way, generators improve continuously and detectors chase. Fix: build editorial processes that don't depend on provenance being knowable, because it increasingly isn't.
FAQ
Can Google detect AI content?
Google can certainly detect the statistical fingerprints of scaled, templated output, that's pattern analysis across many pages, not per-document authorship testing. But its published position evaluates helpfulness and quality regardless of production method, so the practical question isn't "will I be caught" but "is this content worth indexing." Different question, different answer.
Are paid detectors meaningfully better than free ones?
They ship nicer dashboards and bolder accuracy claims. The underlying approach, statistical classification of text properties, carries the same failure modes at every price point: false positives on formulaic human writing, collapse under paraphrase, decay against new models. Nothing about payment fixes the physics.
Should I run my content through a detector before publishing?
As a curiosity, fine; as a gate, no. The score doesn't measure accuracy, usefulness, or originality, the things that decide whether the page deserves to exist. A fact-check pass and an editor's read catch what actually damages you.
My hand-written text got flagged as AI. Why?
Because you write in clear, well-organized, conventional prose, which is statistically what models produce. Formulaic structure, standard transitions, and consistent sentence length all lower perplexity, and low perplexity is most of what these classifiers score. It is a damn frustrating irony: the writing habits every style guide teaches are the ones detectors punish.
Is there any legitimate use for an AI detector?
As a low stakes curiosity or a rough triage signal that prompts a human to look closer, it can have a place. As evidence to fail a student, reject a submission, or fire a writer on its own, it cannot carry that weight, because the score measures statistical texture rather than authorship and produces both false positives and easy evasions.
Does Google penalize content flagged by AI detectors?
There is no evidence that Google runs per document AI detection or penalizes content for being machine assisted. Its published position evaluates helpfulness and quality regardless of how the content was produced, so the useful question is whether the page genuinely helps the reader, not what a third party detector scores it.
Claude Vincent is a technical SEO consultant focused on crawlability, rendering, and AI-search visibility. He writes the field guides and case studies at SEO ProCheck, with a bias toward the durable, unglamorous work that decides whether search engines and AI answer engines can actually read and cite a site.
About SEO ProCheck
Technical SEO consulting and GEO strategy with 20 years of enterprise experience. Case studies, resources, and tools for search and AI visibility.
Work With Me
Technical SEO audits, GEO strategy, site migrations, and international SEO. Hourly consulting for teams who need hands-on support, not just reports.







