Keting Media
Back to blog
Artificial Intelligence
8 min read

Nobody Can Guarantee You GEO Results
(And the Evidence Is Pretty Conclusive)

C

Author

Carlos Beuvrin

Published

Jul 2026

Nobody Can Guarantee You GEO Results (And the Evidence Is Pretty Conclusive)

If someone is selling you "guaranteed ranking on ChatGPT," it's worth asking one simple question before you sign anything: guaranteed measured how, and against what baseline?

The answer usually takes the whole pitch apart. Not because GEO is entirely smoke and mirrors, but because the system that the promise of control is built on is, by design, non-deterministic. And that's not a market opinion — it's measured.

First: what is GEO?

GEO stands for Generative Engine Optimization. It's the SEO equivalent for AI systems that answer with text instead of a list of links: ChatGPT, Perplexity, Claude, Google's AI Overviews. The goal is no longer to rank first — it's to get the AI to mention you when someone asks who to hire.

It's a legitimate field, and there's real work to do in it. The problem isn't GEO: it's what gets promised about GEO.

The number that breaks the promise

Rand Fishkin (SparkToro) and Patrick O'Donnell (Gumshoe.ai) ran the most direct experiment possible: 600 volunteers ran 12 identical prompts on ChatGPT, Claude, and Google's AI, about 3,000 times in total.

The result: the probability of getting the same list of brands twice was less than 1 in 100. The probability of getting the same list in the same order was around 1 in 1,000 — about 0.1%.

Same question. Same model. Same day. And the answer practically never repeats. Not even the length is stable: the same query would sometimes return 2 or 3 options and sometimes more than 10.

Now think about what that means: the full ranking shown in a report — "these five brands, in this order" — has about a 1-in-1,000 chance of happening again. That's not a metric. It's a snapshot of a die mid-air.

One nuance matters here, so as not to overstate the case: that 1-in-1,000 figure applies to the full list in the same order, not to whether an individual brand reappears. Your brand showing up again is considerably more likely than that. What collapses isn't presence: it's position.

It's not a bug in the system. It's the system.

The intuitive reaction is to think "fine, it's noise — enough measurements and it averages out." That's partly true, and that's exactly where the deeper problem lies.

A variance-components analysis across 12,933 responses (GPT-5.2, Gemini 3 Flash, and Perplexity, at temperature 0.3) tried to answer exactly this: of all the variation we see in whether a brand appears or not, how much is actually attributable to the brand itself?

The answer is uncomfortable:

Source of varianceWeight
Query language26.5%
Residual (interactions + resampling)69.3%
Model identity1.6%
Brand identity1.5%
Specific prompt1.1%

Brand — the one thing your agency can actually touch — explains about 1.5% of the variance. A single AI response has an intraclass correlation of 0.0146 for discriminating between brands; the authors describe it as "almost no brand-discriminating signal."

In plain terms: one loose query to ChatGPT about your brand carries no useful information. Not little information. Essentially none.

And there's one more detail: breaking down the residual on a stability subset (7,173 responses), 34.8% of the variance comes from resampling within the same prompt — the model choosing a different continuation for an identical input, even at low temperature. That's structural randomness, not a lack of optimization.

What the research says about GEO tactics

This is where things get genuinely interesting. A critical survey of the 2023–2026 GEO literature reviewed what the evidence actually supports.

On the famous "+40% visibility." This is the number that shows up in almost every sales pitch, taken from GEO's foundational paper. The survey clarifies that the 40% applies only to a document that was already included in a fixed context of five documents handed to the generator. It doesn't measure organic discoverability. It doesn't measure traffic. The underlying experiment is real; the commercial claim "GEO increases your visibility by 40%" is classified by the survey as rejected.

On whether "recipes" transfer. In the C-SEO Bench benchmark, only 3 of 54 method-domain combinations showed statistical significance — and none came out positive for Q&A. What works in one vertical typically doesn't work in another.

On optimizing page content. In SAGEO Arena, optimizing the body text alone reduced top-20 presence by 9%, top-10 presence by 16%, and final citation by 6%. Apparent gains evaporated — or reversed — once the full retrieval stages were included.

On keyword stuffing and "authoritative tone." Consistently null or negative results. Fluency and tone effects are weak and unstable.

On how many measurements you need. Source-level Jaccard scores between days drop to 0.34–0.42, which implies a minimum of 7 to 8 repetitions per prompt just to get a starting baseline.

On a measurement bias almost nobody discloses. In 57.8% of repetitions, ChatGPT didn't even trigger a web search. If your tool only analyzes responses that contain citations, you're measuring a biased sample of your own performance.

And the survey's overall conclusion: no technique reviewed demonstrated a stable, longitudinal, cross-platform causal effect on organic discoverability or user behavior. Confidence that "more citations ⇒ more clicks, conversions, or revenue" is classified as very low.

The ground shifts on its own

Even if you landed a result today, it's not yours to keep.

  • Rephrasing changes everything. In a sensitivity test across 30 query pairs and seven models, on Gemini every single pair changed the cited domains when the question was rephrased.
  • A vendor update rewrites the board. When a widely used engine updates its retrieval rules, results shift simultaneously for the entire user base. You had no part in that decision, and nobody warned you.
  • There's no unified ranking. 53% of the domains cited by Google's AI Overviews don't appear in the organic top 10. Optimizing for one doesn't get you the other.
  • Any edge erodes on its own. C-SEO Bench documented decay as adoption increases: congested dynamics that approach zero-sum. When everyone applies the same tactic, the tactic stops being an advantage.

So is it all smoke and mirrors?

No, and this is the point where honesty is more useful than cynicism. There's real signal beneath the noise, and some findings point in an optimistic direction:

  • In the same Fishkin study, certain names appeared in 60%–90% of responses for a given intent. The pattern concentrated in small, consolidated markets; mass-market categories scattered into chaos. In other words: the ordering is random, but set membership can be fairly stable.
  • A Peec AI analysis of 37,804 responses across five engines found that prompt phrasing matters less than the industry assumes: brand visibility stays stable as long as the core intent is preserved.
  • The survey classifies with moderate confidence that extractable evidence — statistics, definitions, verbatim quotes — produces gains under controlled conditions. That's not the same as a guarantee, but it's about as solid as anything on the table.

None of this is a guarantee. It's something better: it's a basis for working with the right expectations.

How to read a GEO proposal

Red flags:

  1. "We guarantee the #1 spot on ChatGPT." The full ranking repeats about 1 time in 1,000. There is no stable #1 to guarantee.
  2. "We'll boost your visibility by 40%." Ask which paper that number comes from and exactly what condition it applies to.
  3. Reports based on one query per prompt. You need a minimum of 7–8 repetitions, and multiple languages where relevant. One query is an anecdote.
  4. Silence on language. Language accounts for 26.5% of the variance — more than model, brand, and prompt combined. A proposal that doesn't mention it hasn't read the evidence.
  5. No stated baseline. Without an initial appearance rate measured with a real method, "we improved it" means nothing.

What a good provider can honestly promise:

  • Measurement with a disclosed methodology: number of repetitions, prompts, languages, engines, dates.
  • Appearance rate (share of voice) as the metric, not position.
  • Work on what's actually within your control: presence and consistency in the sources these systems actually retrieve from, structured data, extractable and citable evidence, and third-party reputation.
  • Confidence intervals, not point promises.

The bottom line

The difference between a serious provider and one selling smoke isn't technical knowledge. Often it's the same knowledge. The difference is what they do with the uncertainty: disclose it or hide it.

A system where the same question returns the same ranking one time in a thousand doesn't allow for position guarantees. It allows for method, honest measurement, and calibrated expectations. Anyone offering you more than that either hasn't read the evidence, or is counting on you not having read it.

Sources