Somewhere in your CMS there’s a directory you forgot existed. Maybe it’s the integrations pages, one for each of 300-plus connectors, each a swapped product name inside the same 180-word shell. Maybe it’s the location pages your agency built in 2023. Maybe it’s six years of blog posts that three different content vendors wrote to the same outline.
Nobody set out to build so many pages. They just…accumulated. Call it permutation rot: replicated content that made sense one page at a time, but it’s turned into a liability at volume.
The Content Integrity Review measures how much of it you have and whether it creates exposure for you under Google’s scaled content abuse policy.
How can poor content integrity hurt your enterprise?
- Google penalizes the domain, not the pages. The hundreds of pages nobody reads can de-rank the few that actually close deals.
- Your good content gets averaged down. A strong blog on a junk-heavy domain inherits the domain’s problem. You’re buying suppression with all that unread content.
- AI search cites somebody else. ChatGPT and Perplexity pick one source per claim. Give them scores of near-identical pages and they’ll take a competitor’s single good one instead.
Crawl budget goes to the junk first. New content gets found late and updates get re-crawled later, so you lose control of campaign timing.- Your analytics start lying. Thousands of dead pages drag on every average you report to your boss or board.
- The bill grows every quarter. The longer you ignore multiplication of pages, the harder and more costly it is to cull them.
- You find out at a bad moment. During a replatform, an acquisition or rebrand, somebody counts URLs and sees there’s a problem.
Google no longer cares who wrote it
The biggest mistake in content marketing right now is treating this as an AI issue.
It isn’t. Google’s scaled content abuse policy is intentionally method-neutral. It doesn’t matter whether the pages came from a language model, a template, an offshore content shop, or someone manually swapping city names in a spreadsheet. The policy targets pages made mainly to manipulate rankings instead of helping users. The issue is their purpose, not origin.
That cuts both ways. A thoughtful AI-assisted draft can be fine. Four hundred human-written pages that differ by one noun aren’t.
That means the AI detector your agency is using is measuring the wrong thing. It answers a question Google isn’t asking, with methods Google hasn’t defined. It may flag your best writer’s clean prose while missing the integrations directory that’s actually putting you at risk.
The problem isn’t one bad page, but the site-wide pattern
That’s what separates a Content Integrity Review from a standard content audit.
If one page underperforms, you lose that page. If a corpus trips scaled content abuse, the whole site can take the hit. Algorithmic suppression and manual actions often work at the domain or section level, so 300 junk pages can pull down the 40 pages your pipeline depends on. Your best case study, pricing page, or real thought-leadership piece can all sit downstream of a URL template decision.
That’s why this has to be measured across the entire content corpus. A page-by-page quality score misses the source of the risk: the relationship between pages. Duplication only exists in the plural.
AI search makes the problem worse. ChatGPT, Perplexity, and AI Overviews still have to choose what to cite. Hundreds of near-identical pages create weak, scattered signals instead of one clear one, so your competitor’s single strong page is more likely to win. You’re not so much penalized as diluted into irrelevance.
What we measure
The review scores six factors against fixed benchmarks and rolls them into a 0–100 Content Integrity Index. The higher the score, the lower the exposure.
- Duplication Density. Finds near-duplicate clusters across the sampled corpus, using MinHash to catch paraphrased pages as well as copied ones.
- Template Saturation. Measures how much of a page is shared structure—navigation, footer, and boilerplate—versus content unique to that page. DOM skeleton fingerprinting keeps redesigns from hiding the pattern.
- Content Substance. Looks past word count and measures concrete detail: numbers, names, specifics, and other signals that separate useful content from filler.
- Originality Signal. Measures unique 7-gram share across the corpus to see whether pages say meaningfully different things or just reshuffle the same language.
- Publication Velocity. Checks variation in publish dates. Editorial teams publish unevenly; programmatic builds tend to drop hundreds of pages at once and then go quiet.
- URL Pattern Sprawl. Analyzes URL structures to distinguish normal archives from runaway template patterns like /integrations/{a}-{b}/.

What you get
We’ll provide a findings report you can share internally, a page-by-page CSV with the measurements, and a JSON evidence file if your team wants to poke around in the raw data. And you’ll get it in days, not weeks.
The crawl plays by robots.txt, samples across your URL structure instead of just taking the first things in the sitemap, and knows when it’s run into a bot wall instead of real pages: If a Cloudflare-protected site gets read the wrong way, it can look like a stack of identical 200-word pages, which is exactly the pattern this review is looking for. A lot of tools would call your security layer a content crisis. This one stops and tells you what actually happened.
And if everything looks fine, the report says that too. You’re not paying for me to invent a problem just so there’s something scary in the deck.
What this is not
This isn’t an AI detector, a plagiarism checker, or a bloated technical audit that buries the key finding on page four.
It measures the four things Google’s policy actually calls out: duplication, substance, originality, and volume.
Five actual companies and the exposure we uncovered
These weren’t clients or cherry-picked examples. They’re a representative set of B2B software companies we scanned to test our tool. Five of 30 showed meaningful scaled-content risk. Here’s what each one stands to lose.
The company whose good content was getting drowned out
A customer data vendor had more than 19,000 near-identical permutation pages on the same domain as a strong blog, docs library, and guides program. Those sections scored in the high 80s and low 90s. The full site scored 36 out of 100.
What that costs them: Their team is producing content that should rank, but it’s being averaged against 19,000 pages that don’t. The good signals get diluted before they reach the domain. They’re paying for content marketing and buying suppression with it.

The company site that’s all programmatic
For one workflow-automation platform, generated pages weren’t just most of the site. They were the site. The scan hit a 400,000-URL ceiling before it finished counting.
What that costs them: There’s no “clean” content to protect. If Google acts on the pattern, everything is exposed, including pricing pages and case studies. There also isn’t much distinctive for AI assistants to cite, which is why competitors with far fewer pages can get cited instead.
The company with 9,500+ pages
A reverse-ETL vendor had 9,583 pages, with four out of five built from integration permutations.
What that costs them: Real pages are outnumbered five to one in their own sitemap. Crawl budget goes to the permutations first, so new content gets discovered slowly and updates are re-crawled late. Marketing loses control of campaign timing.
The company at 48% programmatic pages
A data-movement company had 7,037 pages, about half of them programmatic.
What that costs them: This is still salvageable. Because half the corpus is real, consolidation and canonical work could improve the score without touching pages buyers value. That window closes if the programmatic share of pages grows.
The company that should be embarassed
A cloud data platform had 7,038 post-conversion thank-you pages submitted to Google for indexing.
What that costs them: Nothing dramatic…yet. But it shows nobody has reviewed the sitemap in years, which is usually fine until suddenly it isn’t.
An entire diagnostic suite
The Content Integrity Review runs as a scoring module inside our AI Visibility Audit, alongside retrieval and citation analysis. Content corpus risk and AI search visibility are the same problem seen from opposite directions, so they’re more useful together.
It also pairs with the Content Effectiveness Audit when the question is what to do with the pages you’ve found. Integrity shows what’s exposed. Effectiveness shows what’s worth saving.
Run it before a migration or redesign, after an acquisition adds another content library to your domain, or whenever someone suggests generating a new page.
Ask us to discover what’s actually in your sitemap. You’ll probably be surprised!