Here are two findings from this month's UK AI Visibility Index that look contradictory until you understand the mechanics of how AI engines generate recommendations.
Bloom & Wild scored 31, the lowest of all 50 audited brands. We asked Gemini "best Online Flower Delivery brands in the UK." Bloom & Wild appeared at the earliest position of any brand in the entire dataset, 121 characters into the response, ahead of competitors with significantly stronger structural readiness scores.
Hotel Chocolat scored 55, comfortably above the bottom of the Index. We asked Gemini "best Luxury Chocolate Retailer in the UK." Hotel Chocolat did not appear in any of the three runs. Gemini answered with William Curley, Charbonnel et Walker, Pierre Marcolini, and Rococo Chocolates.
Across the 50 brands, the pattern repeats. The Pearson correlation between Index score and recommendation prominence is 0.09. Structural readiness, the thing the AI search discipline has spent the past year teaching brands to measure and optimise, barely predicts whether a brand appears when an AI engine is asked who to recommend.
The explanation for this is not particularly mysterious once you look at what an LLM is actually doing when it answers a question. It is also, importantly, not the same as "GEO is dead" or "structural work doesn't matter." Both of those framings miss what is actually happening, and miss where the work continues to compound.
What an LLM is doing when it recommends a brand
When Gemini answers "best UK online flower delivery brands," it is running two related processes at the same time.
The first is parametric recall. The model has absorbed patterns from a very large training corpus, and certain brands have strong directional associations with certain categories built up through repeated co-occurrence. Bloom & Wild appears in thousands of articles, reviews, gift guides, marketing case studies, news pieces, and editorial references about UK direct-to-consumer flower delivery. When the prompt activates the relevant region of the model's internal representation, Bloom & Wild has high activation. The model has, in effect, learned that "UK online flower delivery" and "Bloom & Wild" belong near each other in latent space.
The second is retrieval-augmented grounding. Gemini does not generate purely from internal recall. It retrieves relevant documents at query time and conditions its output on what those documents say. For a query like "best UK online flower delivery," the retrieved documents are predominantly the same publications, comparisons, and aggregators that have been writing about Bloom & Wild for years. The retrieval surface reinforces the parametric activation.
Schema markup on bloomandwild.com is competing with all of this. So is the configuration of llms.txt. So is the cleanliness of the canonical URL declaration. These are real signals, and they matter for the retrieval step. But they are vanishingly small compared to the cumulative weight of a decade of category-defining coverage in the wider web.
Why Hotel Chocolat is absent
The same explanation runs in reverse for the brands we found missing.
Hotel Chocolat is a well-known UK brand. It has retail stores, strong brand recognition, and respectable structural readiness signals. What it does not have, in the corpus the model was trained on, is dominant association with "luxury chocolate." When the model is asked about UK luxury chocolate, the training data points it toward William Curley, Charbonnel et Walker, Pierre Marcolini, and a handful of others whose category association is "luxury" rather than "premium-accessible." Hotel Chocolat sits in a category-adjacent space, and the model's representation of "luxury chocolate" does not strongly activate it.
This is consistent across the five brands we found absent from all three Gemini runs. LookFantastic competes in beauty retail, but the model's "premium beauty retail" representation activates Harrods Beauty and Selfridges first. Oh Polly competes in fashion, but "women's fashion apparel" activates M&S and ASOS first. Very competes in online retail, but "online retailers in the UK" activates Amazon first. Video Games Chronicle competes in games journalism, but "video game journalism in the UK" activates Eurogamer and Rock Paper Shotgun first.
What this means for the discipline argument
The argument this month about whether GEO is a real discipline has been argued largely from first principles rather than from data. The data, at least for established UK consumer brands, says something that neither camp has fully articulated.
The SEO community is partly right that for Google's AI features specifically, the optimisation work is downstream of standard SEO, because Google's AI features are built on its existing search index. They are also partly right that a meaningful subset of GEO tools are measuring outputs (recommendation snapshots) without measuring the inputs that drive them (structural readiness, training data depth, retrieval surface composition).
The newer wave of tools, including our own UK Index, are measuring real signals. But the assumption that a structural readiness score should directly predict recommendation outcomes in mature consumer categories is, on the evidence, not supported. The score predicts the structural floor. The recommendation outcomes depend additionally on factors no single score captures: depth of category association in training data, composition of the retrieval surface for relevant queries, and the cumulative editorial signal that has accrued to the brand over time.
This is not a failure of structural readiness measurement. It is a clarification of what structural readiness predicts and what it does not.
Where structural readiness still does meaningful work
Three situations, in order of immediate commercial relevance.
The first is new and emerging categories. Categories that have formed in the past few years have not had time to develop strong frequency-of-association patterns in training data. There is no dominant association between "vertical AI for legal practice" and any particular brand. There is no settled answer to "best UK carbon-conscious haircare brand." In these categories, the retrieval surface and the structural signals it grounds on become the primary determinant of recommendation. The training data has no strong prior; what the model retrieves at query time wins. Brands competing in still-forming categories that build structural readiness now are competing on the dimension that actually decides the answer.
The second is retrieval-grounded queries even in mature categories. AI engines do not respond purely from parametric recall. The more an engine grounds its answer in retrieved sources, the more the document surface matters, and the more structural readiness affects which sources get pulled. Brands that are easy to retrieve, parse, and cite get more weight in the grounded portion of the response. The parametric portion will still favour the long-established names. But the grounded portion is where new entrants and challengers can actually move the dial.
The third is the compounding flywheel into future model generations. Each new generation of foundation models is trained partly on content from the period immediately preceding it. Content shaped by today's AI search era, including the citations, references, and editorial coverage that structurally-ready brands attract, becomes part of the training corpus for the next generation. The brands building structural readiness now are not just shaping retrieval today. They are shaping the parametric layer of models 2-3 years out. The flywheel runs slowly. It runs.
The four positions, with brands
Mapping the two dimensions together produces a usable framework for thinking about where any given brand sits.
Strong on structural readiness, strong on category association. Tesco at 73, Wise at 71, Bentley at 69, Just Eat at 65, ITV at 63, Range Rover at 61. These brands have the foundations and have the cumulative editorial signal. They appear in their category responses and they appear early. They are the most difficult to displace.
Strong on structural readiness, weaker on category association. Specsavers at 66, Barclays at 66, Bupa at 65, HSBC at 63, M&S at 59, Sky at 59, DFS at 58, PureGym at 58, John Lewis at 57, Dyson at 57, Telegraph at 60. These brands have invested in the foundations. They appear in their category responses, but often later in the list, after brands with deeper category association. The structural work is doing its job, but the cumulative signal it has built so far is not yet strong enough to move them to the front of the answer.
Weaker on structural readiness, strong on category association. Bloom & Wild at 31, Boots at 45, Wickes at 47, Holland & Barrett at 49, Superdrug at 49, Monzo at 51, Next at 51, Boohoo at 51, Citymapper at 52, BBC at 53, Sports Direct at 53, Costa at 56, Greggs at 56. These brands lead their category recommendations on the strength of years of cultural reference. The structural foundations are thinner than the recommendations suggest, which means the position is fragile to category crowding and to shifts in how engines weight signals over time.
Weaker on both. Very at 42, Video Games Chronicle at 43, British Airways at 43, MedExpress at 45, McLaren at 46, Hotel Chocolat at 55, Dunelm at 53, Charlotte Tilbury at 54. Some are still mentioned in their category responses by virtue of residual recognition. Others are absent entirely. This is the position where structural readiness work has the highest immediate return, because the cumulative signal advantage of competitors is not yet so overwhelming that infrastructure cannot close the gap.
The selection bias worth naming
One caveat worth stating directly. The 50 brands in our UK Index were selected because they are prominent UK consumer brands. This makes the dataset systematically biased toward brands with substantial cumulative category association in training data. The findings here describe what AI recommendation looks like among brands that have already cleared that bar. They should not be extrapolated to small businesses, B2B brands, or brands without strong cultural footprint, where the structural floor matters much more because there is no strong parametric signal to overcome.
This is also why the explanation here is not an argument against structural readiness work. It is an argument for understanding what that work does today (shapes the retrieval surface, affects grounded portions of responses), what it does not do today (overcome decades of training-data association in mature categories), and what it does over multiple model generations (shapes the corpus that future parametric layers learn from).
What comes next
The conversation about AI search visibility will mature over the next twelve months in three directions.
First, the structural-readiness playbook will become standardised. The signals AI engines verify will converge across engines, and closing readiness gaps will become well-known practice rather than contested discipline. Most of what is currently called GEO will simply become how technical SEO is done in an AI-search era.
Second, the measurement of category association will become a separate discipline. There is no widely-used tool today for quantifying a brand's depth of association in AI training data and retrieval surfaces. The closest proxies are observational, share of voice across multi-engine prompt panels and citation density across authoritative sources. Whichever tool builds the rigorous version of this first will define the second layer of AI visibility measurement.
Third, the brands that have invested in both dimensions will pull away. The compounding effect of structural readiness and category association is multiplicative, not additive. Brands with both fire on every mechanism the engines use to generate recommendations. Brands with only one are exposed in different ways: the structurally-strong brand is waiting for its category association to catch up, the culturally-strong brand is depending on a cushion that thins as categories crowd and engines evolve.
The UK AI Visibility Index will continue measuring structural readiness because that is what can be measured deterministically and reproducibly. The recommendation outcomes that emerge from readiness plus category association are what we will continue testing per category, per engine, per quarter. Both layers matter. The discipline argument resolves only by holding both at once, and by being honest about which mechanism is doing the work in any specific case.