When AI Has Nothing on Your Company, It Describes Someone Else
/ 8 min read
Summary
The mechanism is well documented, and I have leaned on Mallen and colleagues enough times in this newsletter that regular readers. The practical question is what this changes for SEO, content quality, and AI search visibility.
Most of us have a blind spot when it comes to how AI perceives our businesses. We assume that if our website is clear, our LinkedIn is updated, and our documentation is thorough, the AI will simply reflect that truth. But there is a dangerous gap between what you publish and what a model actually knows.
The real risk isn't that an AI will say "I don't know" when asked about your company. The risk is that it will answer with absolute confidence, using information that belongs to someone else. It doesn't hedge or warn the user that it is guessing. It simply substitutes your identity with the nearest available pattern. A useful companion note is What AI Says About Your Locations, because it looks at a nearby part of the same system. The same pattern also shows up in AI Recommendation Sets Leave Some Brands Out, where the practical question is how the signal becomes visible.
The Invisible Failure of Content Audits
If you are running a marketing or SEO team, your primary tool for quality control is likely the content audit. You look at your pages, check for accuracy, ensure the schema is valid, and verify that your coverage is competitive. When the audit comes back clean, you feel secure.
The problem is that this failure lives in the void. A content audit only inspects what exists. It cannot tell you what is missing from the model's internal weights. You can have a perfect website and a flawless content strategy, yet the AI will still tell a potential customer something about your company that no page of yours supports.
This is a critical distinction for any business leader. You are auditing your inputs, but the failure is happening at the output. Because you don't own the model's weights, you are essentially blind to the substitution until a user asks a question you will never see in your own analytics. This connects with SEO. Don’t Let Claude Do SEO. when the same signal needs a clearer operating decision.
Expert Interpretation: The tradeoff here is between "perceived control" and "actual visibility." We rely on audits because they provide a measurable metric of success. However, the decision to rely solely on internal audits is a mistake because it ignores the "black box" of the LLM. You must shift your focus from auditing your site to auditing the AI's output.
Why AI Fills Gaps Instead of Admitting Ignorance
AI models are not designed to be honest about their limitations; they are designed to be probabilistic. They struggle with "tail" knowledge, the specific, less popular facts that apply to mid market manufacturers or regional service firms. While scaling models makes them better at recalling famous facts, it does very little for the niche details of a specific company.
Research, such as the 2021 work by Longpre and colleagues on knowledge conflicts, shows that models often lean on memorized information rather than the data immediately in front of them. Instead of abstaining from an answer, the model produces a response based on whatever is strongest in its weights.
It is important to separate this from general "hallucinations." A standard hallucination is often random, a fake statistic or a made up case study. Substitution is systematic. It is a directional error based on the shape of the available evidence. If the model has very little data on you but plenty of data on your biggest competitor, it will simply project the competitor's traits onto you.
Expert Interpretation: This matters because it proves that "more data" isn't always the cure. If the model's weights are already biased toward a dominant player in your industry, the model isn't just missing a fact; it is actively substituting a pattern. The decision to make is whether to fight this with more content or by diversifying where your brand is mentioned across the web.
The Danger of Confident Confabulation
To see how deep this goes, look at how this behavior manifests even in professional writing. There are instances of vendor articles warning brands about AI hallucinations that actually use AI generated hallucinations to make their point. These articles might quote a respected academic or a director of a research center, but the quotes don't exist in any paper, talk, or interview.
They might link to the homepage of a prestigious institution like Stanford HAI or reference reports from Gartner and Nielsen, but the specific claims attributed to those reports are fabricated. The AI creates a "plausible sounding" source because it knows what a high authority report should sound like, even if it doesn't have the actual data.
When this happens, the mechanism has moved from the machine to the published page. It shows that the confidence of the AI can override the factual accuracy of the content, leading humans to publish and distribute substitutions as if they were truths.
Expert Interpretation: This highlights a massive trust gap. The tradeoff is speed versus verification. When we use AI to draft "expert" content, we often trade rigorous fact checking for a professional tone. The decision here is to implement a "zero trust" policy for any attributed quote or statistic generated by AI, regardless of how prestigious the source seems.
Four Ways AI Substitutes Your Identity
Substitution is hard to spot because it doesn't always look like a "lie." It usually takes one of four shapes, and each is typically misdiagnosed as a different problem.
1. Silent Analogy
The model finds the nearest well documented neighbor and describes them as you. If your competitor has a very public pricing model and yours is private, the AI may simply state your pricing as if it were the competitor's. From the model's perspective, it is simply providing the most probable description of a company in your category.
2. Staleness as Currency
The model remembers a version of your company from three years ago and presents it in the present tense. It might list a product you discontinued or an executive who left the company long ago. Because there is no "expiry date" on parametric memory, the AI treats old data as current truth.
3. Thin Evidence as Consensus
If there is only one obscure blog post about your company, the AI may present the claims in that post with the same confidence as if they were backed by forty independent sources. The model doesn't signal the volume of evidence it is using, so a fringe opinion becomes a corporate fact.
4. Category Knowledge Applied Specifically
The model knows a great deal about your industry but very little about your specific firm. It fills the gaps by applying general industry standards to your business, effectively erasing your unique value proposition and replacing it with a "category average."
Expert Interpretation: Each of these shapes represents a different type of brand erosion. Silent analogy is a competitive risk; staleness is a credibility risk; thin evidence is a reputation risk; and category application is a differentiation risk. You must decide which of these risks is most damaging to your specific business model.
Why Publishing More Content Isn't the Fix
The instinctive reaction to these failures is to produce more content. The logic is simple: if the AI is missing information, we will provide it. We will write more pages, cover more keywords, and hope that retrieval augmented generation (RAG) fixes the gaps in the weights.
However, this instinct is often misplaced. Research by Sciavolino and colleagues on entity centric questions suggests that retrieval systems often suffer from the same popularity bias as the models themselves. Dense retrievers tend to perform poorly on "entity rich" questions for less common entities.
In other words, the same gradient that makes you "thin" in the model's weights also makes you "thin" in the retrieval process. If you aren't a common entity, the system is less likely to retrieve your specific, new content and more likely to fall back on the dominant patterns it already knows.
Expert Interpretation: This is the most uncomfortable realization for SEOs. The tradeoff is between "quantity of content" and "authority of presence." If retrieval is just another application of popularity bias, then simply adding pages to your own site is a low use move. The decision should be to move toward building external mentions and third party validation that increases your "entity weight" across the broader web.
Identifying Substitution Through Output Monitoring
The only way to find substitution is to stop inspecting your inputs and start watching the outputs. This is a difficult shift because you do not own the interface where these conversations happen. You cannot sample every possible query a user might ask.
You aren't looking for "inaccuracy" in the general sense. You are looking for claims that have no traceable origin in anything you have published. Look for:
Capabilities you do not possess, described in detail. Implementation timelines you have never quoted. Competitor comparisons based on terms you never used.
Crucially, asking a model "Describe my company" is the least useful test. The model will likely pull from your homepage, which is the most obvious piece of evidence. To find substitution, you have to ask the model to perform a task or make a comparison where it is forced to reach deeper into its weights.
Expert Interpretation: This requires a move from "keyword tracking" to "entity tracking." The tradeoff is that this process is manual and non scalable compared to traditional SEO tools. However, the decision to invest time in "adversarial prompting", trying to trick the AI into substituting your brand, is the only way to know where your brand's digital identity is actually leaking.
Comments
Comments are reviewed before they are published. Links are not allowed inside comments.