Gemini local citations overlap only about 40% of the time when you repeat a query, or so goes the sentence carrying Steady Demand’s Grounding Drift study through the trade coverage we read this week, and it is almost right. The 40% is real, it is the study’s most quotable number, and it comes from a local-SEO agency’s dataset of 14,472 citations across 1,487 Gemini queries, 50 US metros and 10 local-service categories. But the study’s own pages say the figure is a cross-phrasing agreement rate, not an ask-the-same-words-twice rate. Those are two different measurements, and the vendor ran both.
The distinction matters because run-to-run consistency is becoming a budget argument. If an AI answer engine names a different plumber every time a homeowner asks, the case for optimising toward it changes, and so does the case for measuring it. A number that describes how three wordings of one question agree with each other does not tell you how stable a single wording is. The study contains a separate, much smaller experiment that does, and that experiment returns a range rather than a point.
So this is a methodology read, not a rebuttal. Steady Demand is a local-SEO agency publishing marketing research, which is the first caveat and travels with every number below. It is also unusually transparent: it names its formulas, runs a non-AI control, audits its own taxonomy and lowers one of its own headline figures. The job here is to separate what the published method can support from what the headline implies, using the same lens we applied to a different citation-decay study three days ago.
- 01The headline 40% is a cross-phrasing number.It compares three differently worded queries per metro-and-category pair, across the groups the vendor counts as complete — it says 498 of 500, which does not quite square with the 13 queries it says failed. The vendor’s own Grounding Drift page sets out that cross-phrasing scope in its note. It is not a measurement of the same words asked twice.
- 02The literal repeat test is separate, small, and a range.Six fixed New York queries, four timed rounds, 72 logged calls. Cited-domain overlap ran from 46.3% same-day back-to-back down to 26.5% when the next day’s own repeats were compared with each other. The vendor calls 40–46% an observed region, not a constant.
- 03The control is the strongest evidence, and the coverage we read skipped it.On the same 500 metro-and-category combinations and the same timed cadence, Google’s own local pack returned the same top listing 90.2% of the time; Gemini returned the same recommended business 7.9% of the time. That isolates the instability to generative synthesis.
- 04The ~60% own-site share has three values in circulation.61.1% before the vendor’s own taxonomy audit, 58.7% after it on the methodology page, 59.9% on the live dashboard. The pages do not say which is current. Quote it as roughly 60% with that caveat, or not at all.
- 05Agency-run, not independent, and scoped to one surface.Steady Demand self-describes as a GBP, Local Services Ads and local-SEO agency. Collection was stateless Gemini API calls with city text in the query and no geolocation parameter. The metro list and raw data are not published on the four study pages we read, and we found no peer review or replication of it.
01 — The StudyWhat was measured, by whom, and how.
Start with the funder, because every number inherits it. The study is published by Steady Demand under the title What AI Actually Cites, with a companion methodology page, a Grounding Drift deep-dive and an interactive dashboard. The page titles of those study pages describe the company as a Google Business Profile and Local Services Ads management and local SEO agency, and the named author, co-founder Ben Fisher, is presented as a Google Business Profile Diamond Product Expert. We found no funding or sponsorship statement on any of the four pages. This is vendor marketing research by a firm that sells local visibility services, which is not a reason to dismiss it and is a reason to read it carefully.
The design, as disclosed: the 50 largest US metropolitan statistical areas by Census population, ten local-service categories, and three fixed phrasing templates per metro-and-category pair. The templates are best {vertical} in {city}, {state}, top rated {vertical} near {city}, {state} and who is a good {vertical} in {city}. That is a 1,500-query matrix (50 × 10 × 3), of which 1,487 completed; the vendor says 13 failed every retry against the Gemini API and were skipped rather than sampled out. It also says 498 of the 500 metro-and-category groups have all three phrasings. Those two statements do not quite reconcile: 13 missing queries have to fall in at least five different groups, so at most 495 groups can hold a complete set of three (our arithmetic: 13 ÷ 3 rounds up to 5). We report both as the vendor states them, with that gap flagged. Collection ran on July 27–28, 2026.
The ten categories, named on the methodology page though not in the coverage we read: Plumber, Roofer, HVAC Contractor, Electrician, Locksmith, Pest Control, House Cleaning, Personal Injury Lawyer, Dentist and Auto Repair Shop. The 50 metros are described by rule (largest first, each represented by its principal city and state) but are not enumerated as a list on any page we checked, so we do not infer them here.
citations analysed
From 1,487 completed Gemini queries, ten categories, 50 metros. Counted at the domain level from the grounding metadata the Gemini API returns, not scraped from a browser. Agency-stated.
per metro and category
Three fixed templates, each run once per pair inside the collection window. The headline 40% compares these three wordings with each other. No identical repeats exist in the main dataset.
not the consumer app
gemini-flash-latest via the Gemini API with the Google Search grounding tool. Stateless, no logged-in session, location passed as city and state text inside the query only. The vendor flags that API and app retrieval can differ.
Two disclosures deserve credit before any critique. First, the citation definition is precise: the API returns grounding_chunks (the cited source plus a Google-hosted redirect URI), grounding_supports (which sentence a citation backs) and web_search_queries (the literal sub-searches Gemini generated). Domains are aggregated from the title field of each chunk rather than the redirect wrapper in uri, which the vendor says it confirmed against live responses before collecting anything. Second, the overlap maths is named: Jaccard similarity on cited-domain sets, computed pairwise and averaged, in a “loose” full-set form and a “strict” does-the-top-recommendation-match form. Both are specified precisely enough to be checked against the API’s own output rather than taken on trust.
02 — The CorrectionTwo different measurements wearing one number.
The study’s homepage callout reads: “Ask Gemini the exact same question, word for word, twice in a row — and the cited sources still only match about 40% of the time. We call this Grounding Drift.” The martech.org write-up that carried the study on August 20, under Danny Goodwin’s byline, restates the ask-twice framing without qualification, and without noting that the publisher is a local-SEO agency. That is also the framing most likely to travel, because it is the one the study’s own homepage leads with.
The Grounding Drift deep-dive says something narrower. Its scope note states that the 40% is the cross-phrasing agreement figure from the full 1,487-query, 50-metro dataset. In that dataset, each metro-and-category pair has three differently worded queries, each run once. The 40% is how much the cited-domain sets of those three wordings agree with each other, averaged over the groups that hold all three phrasings — the vendor says 498 of 500, with the caveat from Section 01 attached. Nothing in the main dataset asked the same words twice.
“The 40% cross-phrasing agreement figure comes from the full 1,487-query, 50-metro dataset described in the main report and its methodology page.”— Steady Demand, Grounding Drift deep-dive, scope note
The literal ask-twice measurement exists, but it is a different experiment on a different scale, covered in Section 03. The point is not that the vendor hid anything; it disambiguates the two on its own pages. The point is that the homepage teaser and the coverage we read collapse them into a single sentence, and the sentence that survives is the one that describes the smaller experiment using the larger experiment’s number.
Cross-phrasing agreement
How much the cited-domain sets for “best X in city”, “top rated X near city” and “who is a good X in city” agree with each other for the same metro and category. Computes to 40%. Answers: does rewording the question change the sources?
Identical-wording repeat
The same words, unchanged, resubmitted across timed rounds. Overlap ranges from 46.3% back-to-back to 26.5% between the next day’s own repeats. Answers: does asking the same thing again change the sources?
Both are legitimate things to measure, and for a local business both are uncomfortable. But they answer different questions, and a slide that says “identical queries overlap 40%” is citing Measurement A for Measurement B’s question. If you need the ask-twice figure, the honest version is a range with a sample size attached. If you need the 40%, the honest version names the three phrasings.
03 — The Isolate TestSeventy-two calls, four rounds, and a range instead of a point.
The identical-wording experiment, as the deep-dive describes it: six fixed queries, all of the form “best X in New York, NY” for dentist, auto repair, house cleaning, HVAC, plumber and personal injury lawyer, run across four rounds with zero wording change, 72 repeat-test calls logged in total. That works out to three calls per query per round, which matches the vendor’s reference to the next day’s three repeats being compared with one another. It is one metro and six distinct queries, against the main dataset’s 1,487 across 50 metros (our arithmetic: 1,487 ÷ 6 ≈ 248, so the main dataset carries roughly 250 times as many distinct queries; the 72 is a count of calls, not of queries).
The overlap depends on when you ask. The chart below is the vendor’s own Jaccard figures on cited-domain sets, by interval. The bars are agency-stated; we have not re-run the calls.
Identical-wording repeat overlap · Jaccard on cited-domain sets, by interval
Source: Steady Demand, Grounding Drift deep-dive, “The Experiment”. Agency-stated; 6 New York queries, 72 calls.“Treat 40–46% as the region we've observed, not a constant.”— Steady Demand, Grounding Drift deep-dive, “The Experiment”
Two readings of that chart. The generous one is that the vendor’s “40–46%” region is honest about its own spread. The stricter one is that the region excludes the chart’s lowest bar: 26.5% is the overlap between three repeats run on the same day as each other, which is exactly the ask-twice scenario the homepage describes. A single New York day, six queries, and the answer to “ask twice, how much matches?” was closer to a quarter than to 40%. With 72 calls, no one should treat either end as a population estimate. That is the vendor’s own advice, and it applies to the headline too.
The deep-dive also reports where the instability comes from. The search strings Gemini generated internally for the same query (the web_search_queries field) overlapped far less than the resulting citations: 0.011 to 0.056 Jaccard across the four rounds, against 0.265 to 0.463 for the cited domains. The vendor’s reading is that the drift originates in query generation, not in the live web changing between calls. That is a plausible mechanism and a useful one for practitioners, because it means the variance sits inside the model’s retrieval step rather than in your listing.
04 — The ControlThe comparison the coverage we read left out.
The most persuasive part of the study is the part the coverage we read left out. Steady Demand ran a direct non-AI control: the identical 500 metro-and-category combinations, on the same repeated-query cadence (up to five timed rounds, five minutes apart), against Google’s own local pack in parallel. The local-pack data came through a third-party SERP data provider’s endpoint rather than direct scraping, which the vendor notes would breach Google’s terms.
The result, on the “strict” metric of whether the top listing or top recommended business matched round to round: Google’s local pack returned the same top listing 90.2% of the time across 4,466 pairwise round comparisons; Gemini returned the same recommended business 7.9% of the time across 5,000 pairwise comparisons. Both figures are for combinations that completed all five rounds.
Top-result consistency · identical design, Google local pack vs Gemini
Source: Steady Demand methodology page, “Local-pack control test”. Agency-run; 500 combinations, up to 5 rounds, 5 minutes apart.This is what makes the drift finding more than a curiosity. The same cities, the same categories, the same timing, the same underlying Google index: one surface held its top answer nine times in ten and the other held it less than one time in twelve (our arithmetic: 90.2 ÷ 7.9 is a little over eleven). Whatever is unstable is specific to the generative synthesis layer, not to local search data as such. The vendor could have led with this, and the coverage we read should have.
05 — Protocol AuditWhat the method pages disclose, and what they don’t.
A reproducibility claim stands or falls on its protocol, so the table below checks the seven questions we ask of any repeat-query study, plus three openness questions, against the vendor’s own pages. One complication the coverage we read flattens: this is four experiments sharing one report. The main 1,487-query dataset, the six-query isolate test, the local-pack control, and a Gemini-versus-ChatGPT comparison each answer some protocol questions differently, so the last column says which sub-study each answer applies to.
| # | Protocol question | Disclosed? | What the pages say | Applies to |
|---|---|---|---|---|
| Repetition protocol (rows 1–2) | ||||
| 1 | How many times was each query repeated? | Partly. | Main dataset: each of the three phrasings was run once per pair; no identical repeats. Isolate test: four rounds, 72 calls across six queries. Local-pack control and the Gemini side of it: up to five timed rounds. | Main: once each. Isolate: 4 rounds. Control: up to 5 rounds. |
| 2 | What was the interval between repeats? | Partly. | Main dataset: not stated beyond “within the July 27–28 window”; no inter-query spacing given. Isolate test: same day back-to-back, same day +3.5 hours, next day (+19–20 hours). Control: five minutes apart. | Main: not disclosed. Isolate and control: disclosed. |
| Session and location (rows 3–4) | ||||
| 3 | Logged in or logged out; what session state? | Disclosed by design. | Stateless server-side calls to the Gemini API (model gemini-flash-latest, Google Search grounding tool). No browser, no consumer account, so the logged-in question does not apply in the browser sense. The vendor’s known-limitations section warns that API behaviour may not equal consumer-app behaviour because retrieval logic can differ by surface. | All Gemini sub-studies. The ChatGPT comparison used a third-party scraper endpoint instead. |
| 4 | How was location simulated? | Disclosed. | City and state as literal text inside the query string only (for example “best plumber in Phoenix, AZ”). No separate geolocation parameter passed. This is a real scope limit the vendor states: it does not test how a geolocated consumer session behaves. | Main, isolate (New York only) and control. |
| Measurement definitions (rows 5–7) | ||||
| 5 | What counts as a citation? | Disclosed. | A grounding chunk returned by the API, aggregated at the domain level from its title field because the uri field is a Google redirect wrapper. Confirmed against live responses before collection, per the vendor. | All Gemini sub-studies. |
| 6 | How is overlap or consistency computed? | Disclosed. | Jaccard similarity on cited-domain sets, pairwise then averaged (“loose”). A “strict” variant asks whether the single top recommendation matched. The 40% and the 26.5–46.3% range are loose Jaccard; the 90.2% and 7.9% are strict top-match. | All. Note the two metrics do not share a scale. |
| 7 | When was the data collected? | Partly. | Main dataset: July 27–28, 2026. ChatGPT comparison: July 30, 2026. Isolate test: the intervals are stated but we did not find a calendar date for the run on the pages we read. | Main and ChatGPT: disclosed. Isolate: intervals only. |
| Openness (rows 8–10) | ||||
| 8 | Is the 50-metro list published? | Not disclosed. | Described by rule (top 50 MSAs by Census population, largest first, each as principal city and state) but never enumerated on the pages we checked. | Main and control. |
| 9 | Is raw data downloadable? | Not disclosed. | No CSV, dataset or API access is offered. The methodology page mentions an internal citations file only while describing how it matched Gemini and ChatGPT records, not as a download. | All. |
| 10 | Any peer review, audit or replication? | None found. | The vendor cites prior work by SISTRIX and a SIGIR 2026 paper as corroborating context. Neither audits this study’s numbers, and we have not independently verified either. | All. |
Read as a whole, the table is more favourable to the study than the headline treatment deserves. Seven of ten questions are answered in full or in useful part, including the two that decide whether anyone could repeat it, the citation definition and the overlap formula. The gaps are the familiar ones for vendor-published benchmarks: no list of the sampled units, no raw data, no external check. What the table cannot fix is the row that matters most for the headline, row 1: the main dataset contains no identical repeats at all, so it cannot by itself support an ask-twice claim of any percentage.
06 — Share of CitationsThe ~60% with three values.
The study’s other headline is share of voice: roughly 60% of Gemini’s local citations point to the business’s own website, more than directories, review sites and forums combined. That directional finding looks solid. The precise figure does not hold still, and the vendor is the one who says so.
The methodology page describes a taxonomy audit of the fifteen highest-frequency domains that had been defaulting to the “business’s own site” bucket. Nine of the fifteen were misclassified: two city magazines, two lead-generation sites, a review-management vendor, two membership directories and others. After correcting the rules, the own-site share moved from 61.1% to 58.7% “at the time of this audit”. The report page rounds to 60%. The interactive dashboard displays 59.9%, which is 8,662 of 14,472 citations (our arithmetic: 8,662 ÷ 14,472 = 59.85%, so the rounding checks). The pages do not say which of 58.7% and 59.9% is current, or whether the dashboard was rebuilt after the audit.
of Gemini local citations
Three agency-stated values: 61.1% pre-audit, 58.7% post-audit on the methodology page, 59.9% on the dashboard. Directionally robust; do not quote a decimal the vendor’s pages do not agree on.
second-largest single source
Query-level bootstrap range 12.9–14.5% across 2,000 resamples. One domain out-cites the entire local-directory category. Agency-stated; Gemini only.
Angi, Thumbtack, HomeAdvisor and peers
Bootstrap range 9.8–10.9%. The vendor reports the Reddit-minus-directory gap at 2.3–4.5 points and says it never crossed zero in 2,000 query-level resamples, so the ordering is stable even if the decimals move.
Two things are worth crediting here. The vendor caught its own classification error and published the lower number, which is the opposite of the usual incentive. And the confidence intervals are bootstrapped at the query level, not the citation level, because citations arrive in clumps of roughly eight to ten per query and are not independent. That is the right call, and the reason for it is stated on the page rather than left implicit. The vendor even names a selection bias in its own significance test on phrasing volatility by vertical, noting that the pair it compared was chosen after seeing which two of the ten categories were the extremes. Naming it does not remove it, but a reader can only discount what has been disclosed.
What the Reddit figure does not tell you is anything about ChatGPT. The same report runs the 1,487-query design through ChatGPT’s search mode via a third-party endpoint, collected July 30, and finds only 8.3% average Jaccard overlap between the two engines’ cited domains and the same top business named 4.2% of the time. That is engine-to-engine disagreement, a third kind of measurement. The vendor also notes an August 20 re-collection of the ChatGPT side after reports that ChatGPT Search had sharply cut Reddit citations, with Reddit’s share of ChatGPT citations falling from 41.7% to 0.0% between the two pulls. Alongside it the vendor reports a same-day 40-query Gemini spot-check that returned a 6.7% Reddit share and calls that “consistent” with its published 13.7% baseline. Read plainly, 6.7% is less than half of 13.7%; with 40 queries and no interval published for the spot-check, it is too small to confirm or contradict the baseline either way, and “consistent” is the vendor’s characterisation rather than a test result. In any case that update is about ChatGPT; it is not evidence about Gemini’s drift rate.
07 — The VerdictWhat the study can and cannot support.
Here is the study reduced to sentences you can put in a deck, with the version the data supports beside the version that travels.
Rewording the question changes the sources
Supported, with scope: in an agency-run July 2026 sample of 1,487 Gemini API queries across 50 US metros and ten local-service categories, three phrasings of the same question shared cited domains at a 40% average Jaccard overlap. Not supported: “the same question twice matches 40%.”
Identical queries drift
Supported only as a small-sample range: six New York queries, 72 calls, overlap between 26.5% and 46.3% depending on interval. The vendor calls 40–46% an observed region, not a constant. Not supported: any single ask-twice percentage.
Gemini is far less stable than the local pack
Supported and under-quoted: same 500 combinations, same cadence, 90.2% same-top-listing for Google’s local pack against 7.9% same-recommended-business for Gemini, on the strict top-match metric. Not supported: comparing 7.9% to the 40% as if one scale.
AI search is non-deterministic
Not supported by this study alone. It is one model (gemini-flash-latest via API, not the consumer app), one country, ten categories, a two-day window, with no peer review or independent replication that we found. The vendor’s own snapshot caveat applies. Corroborating work it cites measures different things on different timescales.
On the corroboration: the deep-dive cites a SISTRIX study by Johannes Beus, published May 1, 2026, which tracked 82,619 prompts and 1,548,213 snapshots across six countries and three platforms for 17 weeks and reported week-over-week domain churn of 5% for Google AI Overviews, 56% for AI Mode and 74% for ChatGPT Search. It also cites a paper by Grossman and colleagues accepted to ACM SIGIR 2026 on generative search showing materially lower response consistency than traditional search. Both are presented by Steady Demand as context, and the vendor is careful to say its own metric differs from SISTRIX’s in kind and timescale. We have not independently pulled either source; treat them as what the study cites, not what this post confirms.
On Google: we checked the Gemini API’s Grounding with Google Search documentation page specifically for this post. It covers how grounding works, pricing, supported models and how to display citations, and says nothing about whether repeated identical calls return consistent sources, nor about determinism or sampling in the grounding step. That is an absence on one page, not a statement that Google has never addressed it anywhere. It does mean a practitioner has no first-party number to put beside the vendor’s.
And on the obvious comparison: this is not the same finding as the citation-decay study we reviewed on August 17. That one, from a different vendor, measured domain survival over months in a single Australian insurance category on ChatGPT and Google, and its headline number described which cited domains were never cited again. This one measures run-to-run and phrasing-to-phrasing agreement inside a two-day window, on Gemini, across US local services, using Jaccard overlap. Decay over time and drift within a window are different phenomena with different maths, and a chart that puts them on one axis is wrong before it starts.
08 — In PracticeWhat a local business should do with this.
Assume, provisionally, that the direction holds: Gemini’s local answers lean on the business’s own site, Reddit punches above the directories, and which sources get cited varies substantially between asks. Three practical conclusions follow, none of which require believing any particular decimal.
Sample, don’t snapshot
If a single Gemini answer can cite a different set of sources on the next call, a one-off “are we cited?” check is noise. Ask the same prompt several times across a day and report the share of runs that cite you, with the run count. Our share-of-voice framework already reports citation as a percentage across a prompt panel re-run on a fixed cadence, not a one-off check.
Your own site still carries the weight
Across three vendor-stated values the own-site share stays between 58.7% and 61.1%. If Gemini is answering from business websites, the local service pages, service-area content and structured business details on your domain are the asset. That aligns with what still drives visibility in Google’s own local pack, which the control shows is the stable surface.
Date every AI citation number
The vendor warns a repeat run next month may not match. Any AI citation figure in a client report needs the collection date, the surface (API or app) and the model. A number without those three is a number you cannot defend when it moves.
The first conclusion is the one that changes how you measure. Our share-of-voice tracking framework makes a version of this move already: a fixed prompt panel, re-run on a weekly cadence and read as a trend rather than a single snapshot. This study is an argument for pushing the same discipline down to the individual prompt, because a citation that appears in one run out of four is a 25% citation, not a win. The second conclusion is reassuringly boring. Everything we know about what still drives local visibility on Google itself and about what correlates with being cited at all points at the same place this study does: the business’s own pages. If your local-SEO programme is already treating AI answers as one more reader of a strong site rather than a separate channel to be gamed, this study is confirmation, not a pivot. That is how we scope agentic SEO engagements for local and multi-location clients: measure AI citation as a sampled rate, and spend the effort on the pages the engines are reading.
Looking forward, the interesting question is not whether Gemini’s local citations drift but whether the drift shrinks. The vendor’s own mechanism, variance in the generated sub-searches, is the kind of thing a model update could change in either direction, and a snapshot from two days in July cannot tell you the trend. If Steady Demand re-runs the same design in a later window and publishes both, that second run will be worth more than the first; a repeat of the identical experiment is the one thing that would let anyone say whether 40% is a ceiling, a floor or a coincidence.
09 — ConclusionA better study than its headline.
Quote the number with its measurement attached, or don’t quote it.
Steady Demand’s Grounding Drift study is agency marketing research, and it is also one of the more carefully documented pieces of vendor research in this genre: named formulas, a real non-AI control, query-level bootstraps, a published self-correction and a limitations section that anticipates its own critics. The fair criticism is narrow and specific. The 40% is a cross-phrasing agreement rate across three wordings, and the ask-the-same-words-twice result is a separate 72-call experiment that returned a range from 26.5% to 46.3%. The homepage and the coverage we read merge them; the vendor’s own deep-dive keeps them apart.
The number that deserves more circulation is the one the coverage we read passed over: on the identical design, Google’s local pack held its top listing 90.2% of the time and Gemini held its recommended business 7.9% of the time. That is the evidence that the instability lives in generative synthesis, and it is the finding a local business should actually act on, by sampling AI citations as a rate and by keeping the weight on the pages the engines read.
Every figure here is the vendor’s, collected from the Gemini API over two days in late July 2026, on one model, with no independent replication. Quote it with that scope and the study holds up well. Quote it as “ask twice, get 40%” and you are citing the wrong experiment.