Two GEO Experiments Challenge Conventional AI Visibility Advice

Shalin Siriwardhana

Summary

We tracked 15 commercial intent keywords across four platforms: ChatGPT, Claude, Gemini, and Perplexity. Every query was run. The practical question is what this changes for SEO, content quality, and AI search visibility.

A close-up shot of a person's hand holding a smartphone displaying a chat interface with a list of recommended agencies, with a physical notebook and pen resting on a wooden table beside it.

Much of the current advice around generative engine optimization ( GEO ) is based on theory, isolated screenshots, or a single campaign. I wanted measurable results, so I ran two structured experiments back to back, tracked the results manually, and documented what worked, what failed, and what changed between the two tests.

The first experiment, for an existing brand we consulted on, cost thousands of dollars and ran over several months. The second was a 30 day cold start experiment a SaaS link building agency with no measurable AI presence when the test began.

Experiment 1: The consultant test

We tracked 15 commercial intent keywords across four platforms: ChatGPT, Claude, Gemini, and Perplexity. Every query was run manually, with and without a VPN, to account for possible location based differences. Our strategy was to place. The practical read is that brand signals need to be consistent enough for both people and AI systems to form a stable view of the company, its expertise, and its trust signals.

The risk is usually hidden in the execution layer. A page can look fine to a human and still fail for an automated visitor if the form, call to action, rendering path, or confirmation step is not accessible enough for the agent to complete the task.

Listicle placement, PR, and guest posts reinforce each other

In this experiment, listicles and PR appeared to work as one system rather than as separate channels. The listicles that earned citations were usually the same ones being amplified through PR. The listicle introduced the claim, the press. The search implication is whether the section improves the evidence around the page, not simply whether it adds more wording. Clear entities, crawlable structure, internal links, and useful context are what make the topic easier to evaluate.

Source decay is rapid and authority matters

Roughly half of the sources stopped being cited within 30 days. One MEXC placement fell from 29 mentions to 11 week over week, while a Triple Review listicle that had been cited consistently declined without any direct intervention. The. The search implication is whether the section improves the evidence around the page, not simply whether it adds more wording. Clear entities, crawlable structure, internal links, and useful context are what make the topic easier to evaluate.

Capitalization and query type matter

The capitalization of queries appeared to affect which sources were retrieved. Capitalized and lowercase versions returned different citations in three repeated tests, although this finding requires further validation. Query type appeared. The practical question is what this changes in the system: the page structure, the evidence presented, the measurement habit, or the way the topic is connected to related work. The same pattern also shows up in Which Brands Are Vanishing from AI Search?, where the practical question is how the signal becomes visible.

The practical value is in connecting the idea to an observable signal. That means deciding what should be checked, what would prove the issue is real, and where the team should make the smallest useful improvement first.

Exact match keywords and answer placement are important

Exact match keyword targeting still appeared to matter. We ranked for "Best LLM SEO Consultant" but barely appeared for "Best AI SEO Consultant," despite the similar intent. The first phrase had a dedicated listicle, while the second. The search implication is whether the section improves the evidence around the page, not simply whether it adds more wording. Clear entities, crawlable structure, internal links, and useful context are what make the topic easier to evaluate.

Experiment 2: The cold start test

The test involved a SaaS link building agency with no measurable AI presence when the experiment began. During the baseline window from April 30 to May 29, the brand had no measurable presence on any tracked platform. The experiment ran. The practical read is that brand signals need to be consistent enough for both people and AI systems to form a stable view of the company, its expertise, and its trust signals.

Focus on observed citations and earned placements

I learned you want to build the target list from observed citations, not from DR alone. Before beginning outreach, I ran all 15 keywords through every platform, logged the sources that appeared, and ranked them by citation frequency. That. The measurement question is whether this signal changes a decision, not whether it adds another number to a dashboard. Useful reporting connects visibility, engagement, and business outcomes without pretending every AI influenced journey will produce a clean click path. A useful companion note is Practical Way to Measure AI Search Visibility, because it looks at a nearby part of the same system.

Measure citations over time

Time to citation ranged from one to 18 days. Two placements were cited the day after publication, while others took 10 or 11 days. Four placements were live but hadn't been cited by the end of the measurement period. The variation suggests. The measurement question is whether this signal changes a decision, not whether it adds another number to a dashboard. Useful reporting connects visibility, engagement, and business outcomes without pretending every AI influenced journey will produce a clean click path.

Owned listicles are a foundation, not a growth lever

This is where I had to revise my original conclusion. After experiment one, I recommended publishing listicles on an owned site because many top ranking brands appeared to benefit from their own content. Experiment two suggested that an. The practical read is that brand signals need to be consistent enough for both people and AI systems to form a stable view of the company, its expertise, and its trust signals.

The useful check is whether this improves the system behind search performance, not only the words on the page. Internal links, crawlable content, clear entities, current evidence, and a sensible page structure all help the recommendation become easier to trust.

Comparative content is effective

The owned listicle improved when it became less promotional and more comparative. On June 23, we updated it to include our leading competitors instead of presenting the brand alone. Mentions of that source rose from four to 49, a. The practical read is that brand signals need to be consistent enough for both people and AI systems to form a stable view of the company, its expertise, and its trust signals.

What the visibility signal actually changes

What the visibility signal actually changes: two GEO Experiments Challenge Conventional AI Visibility Advice: the Practical Angle should be treated as a visibility signal, not a standalone headline. Introduction Much of the current advice around generative engine optimization ( GEO ) is based on theory, isolated screenshots, or a single campaign. I wanted measurable results, so I ran two structured experiments back to back, tracked the results manually,. This connects with AI Search Visibility when the same signal needs a clearer operating decision.

What the visibility signal actually changes: the practical question is whether the page, brand evidence, and surrounding content make the answer easier to trust. If that support is weak, search systems can still understand the topic but fail to connect it confidently to the brand.

What the visibility signal actually changes: that is why the response should begin with an audit of the evidence already on the site before creating a new asset. The fastest improvement is often a clearer page, a better internal link, or a stronger explanation of why the brand belongs in the answer.

Where the evidence needs to be tested

Where the evidence needs to be tested: a single study or ranking observation should not become a strategy by itself. It should become a diagnostic prompt: which source is being trusted, which query pattern is affected, and which part of the site would make that trust easier to earn?

Where the evidence needs to be tested: that keeps the response grounded. The goal is to improve the evidence chain around the topic rather than publish another summary that repeats what every other page already says.

Where the evidence needs to be tested: the important distinction is between a useful signal and a fashionable talking point. A useful signal changes the brief, the page structure, the linking plan, or the measurement view.

Comments

Comments are reviewed before they are published. Links are not allowed inside comments.

Only your name, optional LinkedIn profile, and comment will be shown.