Avoiding AI Convergence - image of AI volume control
Avoiding AI Convergence - image of AI volume control

Michael SemerBy Michael Semer

What you’ll learn:

  • Why AI-assisted content gets rated better by reviewers and worse by buyers, and what the three studies measuring it actually found
  • The five shows that would have been sanded flat by a moderating pass akin to the 1989 NBC memo that called Seinfeld “weak”
  • Why consensus copy removes the only reason a generative engine would cite your company, and the two questions that tell you whether yours has converged

Emily in Paris is a bad show, if you ask my opinion. So 58 million households watched it in a month anyway. It was Netflix’s most-watched comedy of 2020. It picked up Emmy and Golden Globe nominations while French critics were publicly calling it an insult.

The operative words in the last paragraph were “my opinion.” The implications of having forceful POV in an AI content gigaverse is what we’ll dissect a bit here.

What’s the stuff I can’t stand about Emily in Paris? Paris as a duty-free display case. A lead whose cultural ignorance keeps getting filmed as pluck. Conflicts that show up and clear out inside one episode like somebody in the room was fretting about viewers’ blood pressure.

Darren Star, he of Sex in the City provenance, did all of that on purpose. Then said he was surprised anyone was offended, that he was only poking at clichés everybody already knows.

Fine. It’s still a choice, and it’s the kind of choice a large language model would’ve walked back from. That’s why I laugh when you see producers talking about their AI-scripted or visualized productions.

Feed that script to any model and ask it to punch things up. You’ll get a more careful Emily. The French characters get inner lives and arcs. The office gets real stakes. Somebody learns something.

Now it’s a show nobody can attack and nobody can remember, and it doesn’t pull 58 million households, because the things I hate are the thing that made it travel.

That’s convergence. Not robots writing badly. Robots writing to the middle.

Nobody complains about the middle

You’ll miss it, because the individual output looks better.

Doshi and Hauser ran this exercise in Science Advances back in 2024. Writers who got story ideas from an LLM turned in stories that reviewers scored more creative, better written, and more enjoyable than what the unassisted writers managed. Then the same paper counted the thing that, well counts. The AI-assisted stories were more like each other. The writing got narrower. The “shape” of the story became the same across each story.

Padmakumar and He at NYU found the culprit. In their ICLR study, essays co-written with a raw base model lost no diversity at all. Essays co-written with the feedback-tuned model did, with different authors sidling toward each other. The polish and the sameness show up in the same box. Same training step. You don’t get one without the other.

And if your plan was “we’ll just prompt it better,” excuse me while I ruin that idyll: Wenger and Kenett put 22 different models up against 102 humans in PNAS Nexus this March. The model answers clumped together way tighter than the human ones, and they stayed clumped after the researchers controlled for how the answers were structured.

Switching AI vendors won’t save you, either. The models look more like each other than any of them looks like us.

They tried prompting their way out, too. Telling a model to be imaginative and bold nudged its individual score and did nearly nothing about the variability. Cranking the temperature widened things up right until the answers stopped being in comprehensible human language.

So there’s your AI amplitude dial: Consensus on one end, gibberish on the other, nothing usable in between.

Five shows that wouldn’t have survived an AI opinion

  • Seinfeld. The attempt to sand down the edges actually happened in 1989…and somebody kept the memo. NBC ran the pilot past 400 test viewers who called it “weak”, which an executive later downgraded to a “weak weak”. No part of the audience wanted to watch it again. Nobody liked the supporting characters. Viewers felt Jerry needed a better backup ensemble, still my favorite sentence ever written about Jason Alexander, Julia Louis-Dreyfus, and Michael Richards. Warren Littlefield called it a dagger to the heart and admitted they were scared to move on something research had rejected that hard. Think about it in AI terms: A 400-person panel is just a low-resolution language model. It gave the notes a model gives: make it warmer, get him likable friends, quit filming the laundromat.
  • The Sopranos. Nine years, and it ends on a cut to black in the middle of a sentence. Nothing that optimizes for anything tolerates an unresolved ending, and “give the audience closure” is a perfectly good note. Take it, and you hand back the twenty years of arguing that have kept the show top-of-mind since that night.
  • Fleabag. The looks at the camera aren’t just a cute conceit. They’re where she goes to not be in the room with others, which is why season two blows the whole thing up by having the priest catch her doing it. Any consensus-driven reviewer flags fourth-wall breaks as a gimmick that holds the audience at arm’s length. Cut it in draft two and the show falls over.
  • Breaking Bad, “Fly.” Mid-season, mid-plot, one hour, two guys, one room, one housefly. Plenty of people hated it in real time. Nothing that’s optimizing for retention greenlights that episode. Cut it and the plot loses nothing, which is precisely why it gets cut.
  • Twin Peaks: The Return, Part 8. Twenty-odd minutes of black-and-white nuclear abstraction, no dialogue, no plot, dropped into a network revival. That’s as far out as this argument goes. No model and no stakeholder review on earth would permit it to pass.

Emily in Paris and those five flunk the same test from opposite directions. Sand down Emily and she gets less annoying. “Polish” the other five and they’re gone. What’s being removed isn’t quality. It’s variance, but variance is the only part anybody remembers.

Your homepage has blandwidth

The B2B version of this hides in plain sight. That’s because nobody hate-watches a homepage.

Converged B2B copy doesn’t come out bad. It comes out just fine. Category name in the H1, three benefit pillars, a logo bar, trusted by leading revenue teams, request a demo. Legal signs off. Brand signs off. The CEO signs off. Then your buyer opens four tabs, reads four versions of the same page from four competitors, can’t tell you apart, and goes with whoever’s cheaper or whoever G2 mentioned first.

Call it blandwidth. The share of your content that’s alive because nobody in the review cycle had a problem with it.

You can measure the drift into AI-driven blandwidth now, by the way. Kobak’s group at Tübingen combed more than 15 million biomedical abstracts and found at least 13.5% of 2024 abstracts had been run through an LLM, hitting 40% in some subfields. The shift in scientific vocabulary was bigger than the one COVID caused. The tells were the abuse of words like delve, underscore, pivotal, realm. Biomedicine is a field where grown adults fight about comma placement, and it drifted that quickly. Your category page has drifted a lot faster.

Differentiation isn’t the costly part for a company. Citation is.

Generative engines pull from sources that tell them something they don’t have. Copy that recites the category consensus is, definitionally, stuff the model already knows, so there’s no reason to go get it and no reason to name you when the answer comes out. Consensus copy doesn’t just fail to stand out. It deletes the only reason ChatGPT would ever mention your company.

You optimized your way out of the citation.

Which is where clarity quits on you. A page can post a perfect 5 on the Clarity lever of a Messaging Friction Audit and still be swappable with anyone else’s, because clarity only tells you the buyer understood the sentence. Nothing on that page tells you if they could name the vendor an hour later.

For Vendavo, we built a campaign around a number people could argue with: A trillion dollars in profit left on the table due to bad pricing. It brought in 80-plus MQLs and more than 1,200 press mentions. Run that through a consensus review cycle and a trillion softens into significant untapped revenue, because a trillion is the kind of number that starts fights.

Starting the fight was the whole distribution plan. Nobody’s forwarding a link to significant untapped revenue.

So what’s the point, Emily?

Convergence isn’t ruining most B2B content. For most teams it’s the reason the content exists in the first place. Two people covering an entire content function will ship something with these tools and nothing without them, and that’s not a tragedy.

For a good chunk of any portfolio, the standard answer is the right answer. Release notes. Integration docs. Comparison tables. Pricing FAQs. Nobody in recorded history has wanted a distinctive password-reset page. Converge all of it, happily, and go home early.

So the test isn’t which tool you used. It’s what job the piece is doing. If your reader came looking for the standard answer, give them the standard answer and don’t get precious about it. If the piece is supposed to explain why anyone should pick you, then sounding like the standard answer is how you lose the deal. That call gets made asset by asset. In practice, though, one person decides it once for everything.

Whoever owns the content calendar builds an AI production workflow under a volume target (a prompt, a template, or a model they like) and that becomes the default for the blog, the case studies, the category pages, and everything else in the queue. Nothing forces a second look because you’re too busy spitting out content. Two quarters later organic is flat, none of the pipeline traces back to any of it, and the prompt still hasn’t been questioned.

Pull up the last thing you published and find the claim holding it up. Could a competitor slap their logo on that claim and ship it untouched? Then it was never yours. It belongs to the category, and the model already has it.

Here’s a harder one: Could anybody reasonably disagree with it? Because if there’s nothing in there a smart person could fight you about, there’s nothing in there worth citing either.

I still think Emily in Paris is bad. Somebody made a call I get to argue with, which already puts it a few rungs above whatever’s on most B2B blogs.

FAQs

What is AI content convergence?

AI content convergence is the tendency of AI-assisted writing to become more similar across different authors, even as each individual piece gets better. A 2024 Science Advances study found that writers using LLM story prompts produced work rated more creative than unassisted writers, while those same stories resembled each other more closely. Individual quality goes up. Collective variety goes down.

Does AI make content worse?

No, and that’s the problem. Reviewers consistently rate AI-assisted writing as better written and more enjoyable than the unassisted version. What it loses is variance. The result is content that passes every review and that a buyer can’t distinguish from three competitors’ versions of the same page.

Can better prompts fix AI homogenization?

Not meaningfully. Wenger and Kenett tested 22 models against 102 human participants and published the results in PNAS Nexus in 2026. Prompting models to be imaginative raised individual originality slightly and barely moved variability between responses. Raising the temperature setting increased variability only as answers degraded into incoherence.

Does switching AI models help?

No. In the same PNAS Nexus study, responses from different models clustered together far more tightly than responses from different humans, even after controlling for how answers were structured. Models resemble each other more than any of them resembles a human writer, so changing vendors moves you within the same narrow band.

Why does AI-written content hurt AI search visibility?

Generative engines cite sources that contain information the model doesn’t already hold. Content that restates category consensus is information the model already has, so there’s no reason to retrieve it and no reason to name the company in the answer. Converged content removes the mechanism by which ChatGPT, Perplexity, or Google’s AI Overviews would cite you at all.

Can you detect AI-written content at scale?

Yes, through vocabulary drift. A 2025 Science Advances analysis of more than 15 million biomedical abstracts found at least 13.5% of 2024 abstracts had been processed with an LLM, rising to 40% in some subfields. The markers were words like delve, underscore, pivotal, and realm, appearing at rates that exceeded the vocabulary shift caused by the COVID-19 pandemic.

Should B2B marketing teams stop using AI for content?

No. For release notes, integration documentation, comparison tables, and pricing FAQs, the consensus answer is the correct answer, and AI production is appropriate. The distinction is per asset: use it where the reader wants the standard answer, and avoid it where sounding standard is how you lose the deal.

How do you tell if your content has converged?

Two tests. First, take the load-bearing claim in your most recent published piece and ask whether a competitor could publish it unchanged under their own logo. Second, ask whether anyone could reasonably disagree with it. A claim that survives both tests belongs to the category, not to you, and there’s nothing in it worth citing.

Spread the word: