The Pre-Publish Quality Gate for AI-Assisted Content

No Comments
The pre-publish quality gate for ai-assisted content

TL;DR

A pre-publish quality gate is a short set of pass or fail checks an AI-assisted draft must clear before it ships, and it is the difference between using AI for drafting and building the scaled-content footprint Google suppresses. This is the five check gate we run on every page of this site, built after watching scaled-content suppression cut impressions by roughly 90 percent on a similar site we tested with.

  • The five checks: the anyone test, the no-AI test, says something new, the first screen, and the recommend test. All five must pass.
  • Google's policy is scaled content abuse: it targets mass production without value, whatever tool produced it.
  • The footprint that gets caught is visible in your own library: repeated heading skeletons, near identical page structures, word counts in the low hundreds.
  • A failing page gets another editing round, not publication. On this site failing pages are enriched, never pruned.

AI made drafting cheap. It did not make publishing cheap. The real cost moved from writing to judging, and most content pipelines never rebuilt the judging step. The result is a specific and well documented failure mode: a site publishes hundreds of polished looking pages that add nothing, holds its breath while they rank, and then loses visibility sitewide when Google's scaled-content systems catch up.

That is not a theoretical risk we read about. We watched it land on a test site similar in profile to this one, and the recovery work there is still running. What follows is the pre-publish gate we now run on every page here, enriched or new: five checks, and a draft must pass all of them. It exists so a recovery never rebuilds the footprint that caused the demotion.

What scaled-content suppression looks like from inside

The damage on the test site was easy to read once it landed. Search impressions fell by roughly 90 percent. The Page indexing report showed 2,780 pages parked in Crawled, currently not indexed, a number that plateaued and stayed there. Hundreds of published posts weighed in between 137 and 364 words.

The more useful finding came from auditing that site's library itself. Dozens of posts shared one identical three section skeleton, with only the topic noun swapped in. Any single page looked acceptable on its own. Together they formed a fingerprint, and at scale is exactly how a classifier reads a site. That repeated skeleton, not any individual thin page, was the footprint.

The timing matched a pattern practitioners keep reporting: programmatic content performs for roughly two months, then falls off a cliff. A practitioner at an AI visibility vendor described the same two month arc across client sites in a recent r/SEO_for_AI discussion about pre-publish checks, the thread that prompted us to write this gate down. That site's Search Console graph is that arc.

Google's name for the policy is scaled content abuse: generating many pages primarily to manipulate rankings rather than to help people, regardless of how the pages were produced. Since the March 2024 update it is enforced, not advisory, and there is no manual action to appeal because the suppression is algorithmic. The gate below is what came out of that recovery, and every page on this site now runs through it.

The gate: five checks, every page, all five

Flowchart of the five check pre-publish quality gate for ai-assisted content: the anyone test, the no-ai test, says something new, the first screen, and the recommend test. All five passing leads to ship it; any failure leads back to editing.
The gate at a glance. A draft that fails any check goes back to editing, not to the publish button.

Run these before any completeness pass, meaning images, tables, FAQ, schema. Substance first, presentation second. Each check has to be demonstrable in the page itself, not asserted by whoever wrote it.

1. The anyone test

Could this page have been written about any site, by anyone? If yes, it is a template wearing your logo. A passing page contains at least one element only you would publish: first-party data, a configuration you actually ran, a failure you can document with dates, or a position you are prepared to defend.

To run it on one page, swap the nouns: replace your brand and topic terms with a competitor's and reread. If the page still reads true, it argues for nobody and belongs to no one.

To run it on a whole library, look for repeated structure instead of reading every page. First export your posts as files: in WordPress that is Tools, then Export, or wp export if you use WP-CLI; any CMS export that produces one file per post works. Unzip the export into a folder, open a terminal next to that folder, and run:

grep -rl "Implementation Considerations" wp-export/ | wc -l

The command counts how many files contain that exact heading; swap the quoted text for any section title you suspect is boilerplate, and run it once per suspect heading. One repeated heading is coincidence. The same three heading skeleton across dozens of posts is the fingerprint we found in the test site's library, and it is what a classifier sees.

2. The no-AI test

Would this page be worth publishing if AI drafting did not exist? The question strips the economics out of the decision. If the honest answer is that the page exists only because generating it cost nothing, then its value is also nothing, and shipping it trades a little traffic now for a sitewide problem later.

This check kills entire content plans, which is the point. Cheap production is a reason to raise the publishing bar, not lower it, because everyone else's production got cheap at the same time yours did.

3. Says something new

Does the page contain at least one insight, example, or number the top ranking pages do not already have? To run it, open the top three results for the target query and list their H2s next to yours. If your outline is a subset of theirs, you wrote a summary, and the index does not need another one.

New does not have to mean a proprietary study. A worked example with real numbers, a config that failed and why, or an honest statement of what cannot be measured all pass. Where consensus genuinely is the whole story, say so on the page and name the open question. Honesty about limits is itself differentiation, and it is the position this site's AI visibility work is built on.

4. The first screen

Is the searcher's question resolved before the first scroll? The answer belongs at the top, in plain language, with the depth underneath for readers who want it. That order serves human skimmers, featured snippets, and AI engines that extract answers, all at once.

Teasing the answer to hold attention fails every one of those audiences, and it is the most common failure mode of AI drafts, which love a long runway. If your first screen is throat clearing, the page fails.

5. The recommend test

Would you send this URL to a colleague who asked the question, over the current number one result? This is the human veto. It cannot be automated, and that is its value: it forces one accountable person to compare the draft against the best thing that already exists, not against a rubric.

If the honest answer is no, you already know what happens next, and it is not clicking publish.

The gate at a glance

#CheckThe questionQuick way to run it
1The anyone testCould anyone have written this about any site?Swap the brand and topic nouns; grep the library for repeated heading skeletons.
2The no-AI testWorth publishing if AI drafting did not exist?Ask it out loud in the content meeting; watch which plans survive.
3Says something newOne insight the top results do not already have?Diff your H2 outline against the top three ranking pages.
4The first screenIs the query resolved before the reader scrolls?Load the page, do not scroll, ask if the question is answered.
5The recommend testWould you send this over the current number one?One named person answers yes or no. No committee.

Running the gate at scale

On a personal blog the gate is just reading with intent. At tens or hundreds of pages a month, including agent-driven pipelines, it has to be operational or it will be skipped. The version that works: the gate is written into the reviewer step, and the reviewer must produce evidence, not a verdict. Name the only-you element for check one. Paste the outline diff for check three. Quote the first screen answer for check four. A reviewer who cannot produce the evidence has not run the gate.

Be honest about what automates. Checks one and three are toolable, check four is a ten second glance, and checks two and five are human judgment by design. That split is a feature: the automatable checks catch the footprint, and the human checks catch the hollowness that formatting hides.

Failing pages get work, not the trash can. Our rule on this site is enrich only: a thin page attached to a real query keeps its URL and gets rebuilt until it passes, because the query is the asset and the filler was never the point. Deletion is reserved for pages with no query and nothing salvageable. Recovery, meanwhile, is measured in the same reports that showed the damage, the not indexed count shrinking and impressions returning section by section, and we report those numbers rather than promising dates, the same way we treat organic search ROI reporting generally.

What the gate does not do

The gate checks substance, and substance is not the whole job. A separate completeness pass covers the rest: a matched image with real alt text, at least one genuinely useful table, an FAQ, schema markup, and internal links that actually resolve. Those matter, but they come second, because a complete hollow page is still hollow.

The gate also does not measure outcomes. Whether AI engines actually surface your pages is its own discipline, with its own traps, and we keep that in one place: Monitoring Your AI Search Visibility. And before quality enters the conversation at all, the page has to be retrievable: most AI crawlers do not execute JavaScript, so a page they cannot read passes no gate that matters to them.

FAQ

Does Google penalize AI-generated content?

No. Google's spam policy targets scaled content abuse: producing many pages primarily to manipulate rankings rather than to help users, no matter what produced them. Human-written filler at scale is just as much in scope as AI output, and AI-assisted pages that genuinely help people are fine. The gate exists because scale without judgment is the risk, not the drafting tool.

How do I know if my site was hit by scaled-content suppression?

The pattern is sitewide, not page level: impressions decline sharply over a few weeks, the Crawled, currently not indexed bucket in Search Console grows and then plateaus, and whole sections sit at zero impressions while a handful of strong pages keep ranking. There is usually no manual action; the suppression is algorithmic, so nothing appears in the Manual Actions report.

Should I delete, noindex, or enrich thin AI pages?

Our approach on this site is enrich first. A thin page attached to a real query is an asset with missing work, and deleting it throws away the query along with the filler. We reserve deletion for pages with no query and no salvageable intent, and we do not mass-noindex as a reflex. Whichever route you choose, the gate applies to the rewritten page before it goes back out.

Can the gate be automated?

Partially. The anyone test and says something new can be assisted by tooling: grep your library for repeated heading skeletons, and diff your outline against the top ranking pages. The first screen check is a ten second glance. The no-AI test and the recommend test are judgment calls, and if nobody in your pipeline is accountable for those two, you have a formatter, not a gate.

How long does recovery from scaled-content suppression take?

There is no fixed timeline, and anyone quoting one is guessing. Recovery shows up in the same reports that showed the damage: the Crawled, currently not indexed count shrinking, and impressions returning section by section as pages are rebuilt. Expect months, report progress with those two numbers, and do not promise dates.

Do AI answer engines cite AI-assisted content?

They cite what they can retrieve and extract, and there is no public evidence that any of them detect and exclude AI-assisted text. In practice the same pages that fail this gate rarely earn citations anyway, because a page that adds nothing gives an answer engine nothing to synthesize. Retrievability comes first, which is a separate check from quality.

Sources

Claude Vincent is a technical SEO consultant focused on crawlability, rendering, and AI-search visibility. He writes the field guides and case studies at SEO ProCheck, with a bias toward the durable, unglamorous work that decides whether search engines and AI answer engines can actually read and cite a site.

    About SEO ProCheck

    Technical SEO consulting and GEO strategy with 20 years of enterprise experience. Case studies, resources, and tools for search and AI visibility.

    Work With Me

    Technical SEO audits, GEO strategy, site migrations, and international SEO. Hourly consulting for teams who need hands-on support, not just reports.

    Subscribe to our newsletter!

    More from our blog