By Michael Semer
What you’ll learn:
- Why AI makes content cheaper to produce but harder to get cited — and why volume was never the thing standing between you and being found.
- What an answer engine actually lifts from a page, and the four traits that make a claim citable: self-contained, answer-shaped, specific, and anchored to you.
- Where the widely quoted “40% visibility gain” holds and where it doesn’t — plus what has to be true before extractability matters at all.
Congratulations. You can generate a quarter’s worth of blog posts before lunch with the help of AI. And, of course, so car your competitor, using the same model and very nearly the same prompt. And the AI that wrote it for you is also the search engine deciding what to cite. And it won’t cite you.
The challenge of churning out drafts has been definitively solved. But it was never the constraint. Think back to the days of content mills and underpaid freelancers. There was plenty of craptent being published.
Today, the cost of production has fallen to roughly zero for the kind of content that many marketers are willing to publish. And that’s driving convergence, though not in a good way.
Language models learn by absorbing vast amounts of text and picking up patterns in how words tend to follow other words. When they write, they’re basically answering, “given everything I’ve seen, what word most likely comes next?”
So they reach for the most common, expected phrasing. That’s their default.
Two factors push every model’s output toward the same bland, undifferentiated center:
- First, they all trained on roughly the same internet, which is already overstuffed with mediocrity. So they’ve absorbed the same common phrasings and arrive at similar answers.
- Second, “most likely next word” is an averaging move, favoring what’s typical over what’s distinctive. The unusual angle, the specific number, the contrarian take… those are statistically rare, so the model tends to bypass or soften them in favor of the safe, familiar version.
Put those together and you get a the Great Vanilla Convergence. Ask ten companies to write about the same topic with the same tool, and you get ten drafts that read like slight rewrites of one another.
Not because anyone copied, but because everyone drew from the same well and asked for the most probable phrasing. All in pursuit of just grinding out more content.
You end up at the statistical center; the average, expected output. Unless you force them somewhere else, it’s where these new tools will naturally land. And that’s the reason why AI drafts are hard for a search engine to cite: there’s nothing in the average that’s uniquely worth you being specifically attributed for putting it out there.
Remember what happens where there’s an ubiquitous supply of undifferentiated products: The product becomes totally commoditized or, worse, relatively worthless.
The buyer is actively ignoring you
During the evaluation phase of the sales cycle, the decisive moment that you, as a seller, is interested in? That happens before you’re ever in the room, so to speak.
Increasingly it happens inside an answer engine. Buyers spend most of their journey in independent research, and 94% of them will rank a preferred vendor before first contact. And that favorite during the preliminaries will capture the sale about 80% of the time.
But even though 89% of purchases now include AI as part of this process, most seller sites fail to address them by designing their content around AI visibility. It’s an AI information gap the engine can’t close, or (more properly) a citation gap. You might have a better story to tell or a better product, but AIs can’t see it.
The net-net is, buyer preferences are forming during an extraction event that you aren’t attending. It’s the moment an AI answer engine (ChatGPT, Perplexity, Google’s AI answers) pulls a specific claim off a seller’s page and repeats it in the answer it gives a buyer, naming where it came from. Which may not be your site.
Drafting and extraction are different acts
Let’s see what drives citation, then. What, exactly, is extraction?
It’s when a generative engine synthesizes and answer and cites the sources it can pull a clean, attributable claim from. What raises the probability of being cited? Statistics, quotations, expertises, contrarian or falsifiable claims. Keyword density and classic SEO signals barely move it, though some AEO/GEO doubters say the opposite (hint: they’re wrong).
No amount of SEO will fix what’s broken if AI-drafted consensus prose has nothing worth lifting, which means it’s structurally un-citable.
That’s why tools like a Messaging Friction Audit or AI Visibility Audit are needed to see if content is going to stand any chance of being extracted and cited, and help a marketer understand what need to be fixed if there’s an issue.
Why an AI draft is un-extractable – by design!
I’m going to beat this one to death, because it’s so important to understand if you’re a digital marketer: AI models regress to the mean and hand every user the same phrasing. But extraction demands the opposite: a specific, ownable claim that isn’t already all over the web.
It’s why executive thought leadership only happens when there’s actual thought and leadership taking place, when content is intrinsically citable because it’s truly authoritative and distinguishing. I
See, it’s important to underststand that whether or not a sentence gets lifted into an AI answer is a property of the sentence itself, and four traits make it liftable:
- Self-containment (it stands on its own without the paragraph around it)
- Answer shape (it reads like a direct reply to a question, not a windup to one)
- Specificity (a real number, name, or claim, not a vague gesture)
- An attribution anchor (something that ties the claim to you, so quoting it means naming you)
But there’s a trap. You can nail all four and still draft a sentence that says exactly what every competitor says. You’ve just uttered a clean, quotable version of an industry consensus. The AI engine may happily cite it, because it’s easy to lift, and you get named.
What you don’t get is authority: you’ve spent real effort making a generic point extractable, yet all it buys you is being one more voice repeating what everyone already knows.
That’s wasted specificity, a citation without the credit for saying anything only you could say.
Understanding the conditions behind AI visibility
The paper that first gave a name to generative engine optimization, GEO: Generative Engine Optimization (Aggarwal et al., Princeton, KDD 2024), reported that its methods raised a source’s visibility in AI answers by up to 40%.
That figure gets quoted everywhere, but often is stripped of the conditions that make it true.
The 40% is a before-and-after comparison. The researchers took a page, rewrote it with statistics, quotations, and cited sources, then measured how much bigger a share of the AI’s answer that page got. The rewrite earned up to 40% more space in the answer than the original version.
Not 40% more traffic, and not a 40% chance of appearing in ChatGPT.
That’s the catch. The researchers gave the engine the top five pages for each question before anything was measured. The test showed how much a rewrite grows your share of the answer once you’re already in that group of five. It says nothing about how to get into the five in the first place.
That gap is the game, though; getting retrieved is the real problem. A critical survey of GEO research found that later work in the field points to relevance and position (whether the engine pulls your page at all, and where it sits) as bigger drivers of the first citation than any wording change. One study found that moving a source higher in the context beat most rewrites outright.
So: get retrieved, then get extracted. Relevance, entity clarity, and corroboration decide if you’re in contention. Extractability decides your share of the answer once you are. A perfectly liftable page that never gets retrieved might never get read.
What extraction actually takes (the human work)
Extraction is one just lever of six for obtaining AI visibility. An AEO/GEO Visibility Audit scores all of them: Machine Accessibility, Structured Data, Entity Definition, Answer Extractability, Corroboration Density, and the Conversion Bridge.
Three are dev tickets. Machine Accessibility asks whether a crawler or an AI agent can render and read the page at all, which comes down to JavaScript gating, robots rules, and load speed. Structured Data asks whether your schema is present and valid, so an engine parses what’s on the page instead of guessing. The Conversion Bridge asks whether the buyer an engine sends you lands somewhere that converts or somewhere that dead-ends. A competent developer can close all three.
The other three can’t be closed by tooling. They need someone to originate something true.
Entity Definition is whether the web states, consistently, what you are and what you do. A model can format that claim once you’ve made it. It can’t decide what you are.
Answer Extractability is whether your page holds a clean claim worth lifting. Structured Data can wrap the block in schema, but the sentence inside still has to say something specific, and specificity is the thing an AI draft removes.
Corroboration Density is whether independent sources vouch for the claim, things like press, reviews, and analyst mentions, so an engine has more than your own say-so to cite. You can earn that. You can’t generate it. Your named results are corroboration. A competitor’s AI-written case study of itself is not.
So the bottleneck isn’t the machine work. It’s the four things an AI draft can’t originate: a proprietary claim, a defined entity, a structured answer block with an actual answer in it, and independent proof someone else wrote. That’s the part of an AI Visibility discipline that requires human involvement.
How to test it on your own site
There are three checks that you, dear reader, can run today to test your AI visibility. Some of them are subjective and very human, and that’s a good thing.
- Strip the logo and author bio from your top post and see if you can tell it from a competitor’s article.
- Tun your buyers’ category queries in ChatGPT and Perplexity and note who gets named.
- Take the block on your page that’s supposed to answer a buyer’s question, and read it for five seconds (about as long as an engine or a skimming buyer gives it). In those five seconds, can you state the specific claim the page is making? If you can’t, neither can an AI, so there’s nothing for it to lift.
FAQs
Can AI write content that gets cited in AI search results?
AI can produce the page. It can’t make an answer engine quote it. Citation depends on whether the page holds a specific, liftable claim, and generic AI drafts are the kind of writing that has nothing to lift.
What is an extraction event?
An extraction event is the moment an AI answer engine pulls a claim off your page and repeats it in the answer it gives a buyer, naming where it came from. Either it finds a clean, quotable line and lifts it, or it finds nothing usable and pulls from a competitor or a review site. This is where the buyer forms an opinion about you, without a click or a visit you can see.
What makes a claim citable by an AI answer engine?
Four traits. It stands on its own without the surrounding paragraph. It reads like a direct answer to a question. It contains a real number, name, or claim rather than a vague gesture. And it’s anchored to you, so quoting it means naming you.
Does the “40% visibility gain” from the GEO study mean my content will show up in ChatGPT?
No. The figure is a before-and-after: a rewritten page earned up to 40% more space in an answer than its original version. In the study, the engine was already handed that page as one of five sources. The number describes your share once you’re in the group. It says nothing about getting in.
Why does AI-generated content struggle to get cited?
Language models reach for the most probable phrasing, which is the phrasing every other model also produces. The result reads like the industry consensus. An engine has no reason to quote one more copy of what every page already says.
What actually gets your content cited?
Two steps, in order. Get retrieved, which depends on relevance, a clearly defined entity, and independent corroboration. Then get extracted, which depends on whether the page holds a claim worth lifting. Writing is cheap. Being the source worth citing is the work.
By Michael Semer