The Problem: Famous, And Invisible
We ran a simple test before the engagement started. We asked several AI assistants about the brand's events: distances, qualification, locations, history. The answers were accurate and detailed, and when we scored them by topic the brand landed between 95 and 100.
Then we asked the questions a prospective athlete actually asks. How do I train for my first long-course race. What should I eat the week before. Am I ready. Which bike. On those, the brand barely appeared. Magazines, coaching platforms, forums, and a competing event organizer answered instead.
That gap is the whole story of AI search for established brands. The models already know who you are. Fame is not the constraint. The constraint is that on any question where the assistant has to go and look, it goes to whichever pages rank and are shaped like answers.
The brand's training and nutrition content was neither. Most of it was years old, structured as long editorial essays, with no structured data, no direct answer near the top, and internal links that pointed at race registration rather than the next question in the athlete's head.
The commercial stakes were not abstract. The brand's two primary conversions were app downloads and race registration, and both start with a person who is deciding whether they are the kind of athlete who does this. If the assistant they ask sends them to a magazine, the magazine gets the relationship. The brand gets the registration only if the athlete finds their way back.
What made this hard to act on was measurement. Nobody on the client side could say, with numbers, how visible the brand was across ChatGPT, Google AI Overviews, AI Mode, and Gemini, or whether a change on the site moved any of them.
Without a baseline you cannot tell the difference between a program that is working slowly and one that is not working at all. And, as we found out later, you cannot tell the difference between four engines failing for four reasons and four engines failing for one.
What We Did And Why
Benchmark First, By Topic, Per Engine
We set up Conductor's AI search tracking against a prompt set built from the brand's real questions, not its keywords. That distinction matters. A keyword list says "marathon training plan." A prompt set says "I have 16 weeks and I have never done this, what should my training look like."
We grouped prompts by topic (training, nutrition, gear, readiness, race logistics, event-specific) and tracked each engine separately: ChatGPT, Google AI Overviews, AI Mode, and Gemini. The output was a visibility score per topic per engine, plus the domains being cited in the brand's place.
That last column changed the plan. When we looked at who was being cited on training and nutrition prompts, it was almost entirely editorial and community sites with answer-shaped pages: a direct answer in the first paragraph, an FAQ, a clear heading per sub-question, and schema. The brand's competitors were not out-writing it. They were out-structuring it.
Make The Pages Machine-Readable
The refresh program (covered in a companion piece on reviving a dead content library) gave us the vehicle. Every article going through the cycle got a scaffolding pass built specifically for AI retrieval, generated by our in-house agentic content system and validated before the client saw it.
The pass does five things.
- It writes a short answer capsule for the top of the article, a direct two-to-four sentence answer to the primary question, so an engine extracting a response does not have to hunt for it.
- It restructures the headings so each maps to a real sub-question from the prompt set, and moves the answer to the first sentence under each heading.
- It builds an FAQ block from the questions Conductor's brief surfaced, answered in the article's own voice rather than generic filler.
- It generates and validates Article and FAQ structured data, so the page declares what it is rather than leaving the engine to infer.
- And it proposes internal links so the training article points to the nutrition article points to the readiness article, giving the engines a cluster to trust rather than an isolated page.
We deliberately did not do the thing most AI visibility advice recommends, which is to stuff the page with the brand name and hope the model notices. The models already knew the brand. What they needed was a reason to retrieve its pages, and retrieval follows structure and ranking, not mentions.
Keep Conductor As The Referee
Once a refreshed page was live, Conductor did two jobs. It scored the page against the brief, which told us whether the restructuring had actually closed the gap with the pages being cited. And it kept tracking the prompt set, so we could see, week by week and engine by engine, whether ownership was moving.
We reviewed that view every week with the client's content and platform leads and used it to decide which topic cluster to refresh next.
The alternative we considered was treating AI visibility as a separate workstream with its own content. We rejected it because it would have doubled the production load and split the authority signals across two sets of pages. Folding the scaffolding into the refresh cycle meant every page we touched for organic search was also being shaped for AI retrieval, at no extra review cost to the client.
What Changed
The number we watched most closely was owned AI Overview results: keywords where a brand page is the source Google's AI Overview cites. For the first three weeks after the refreshed pages went live, it sat at 3. Then it went 36, 66, 99, 195 over the following six weeks.
Ninety-nine keywords with an owned AI Overview, up from three, on pages that had been earning nothing a quarter earlier. One nutrition article on its own is generating more than 200,000 impressions in a 30-day window, largely through AI Overview placements.
AI-referred sessions, meaning visits where the referrer was an AI assistant rather than a search results page, reached 210 in the nine-week window. That is a small number next to the roughly 16,000 organic sessions the same pages earned. It is also a number that was effectively zero before, and it is growing week on week.
The honest edge, and it is a big one. Every one of those 210 AI-referred sessions came from ChatGPT. Zero from Perplexity, zero from Gemini, and the AI Overview gains show up as Google impressions rather than as AI referrals. When the client asked why, our first hypothesis was about content style: ChatGPT favoring accessible writing, other engines favoring denser, more data-heavy sources. That is partly true and mostly beside the point.
The better explanation is architectural. ChatGPT leans heavily on what it learned in training, and a famous brand gets named from memory whether or not its pages rank. Gemini, AI Overviews, AI Mode, and Perplexity are retrieval-grounded: they run a live search and cite what comes back. Three of those four share Google's index, so they succeed or fail as a block.
If the refreshed pages are only now climbing into the top results for preparation queries, the retrieval engines are only now starting to see them. The AI Overview curve, three weeks flat then a sharp climb, is exactly what that lag looks like from the outside.
We expect the referral mix to diversify over the next 30 to 60 days as the pages mature in the index. We do not yet have the data to prove it, and we are saying so.
Two other things we are checking that any team in this position should check: whether the site's robots rules or CDN are blocking the crawlers those engines use (a common and invisible cause of exactly this pattern), and whether the tracking platform's prompt runs are geo-locked to a region where some AI surfaces do not trigger.
What You Can Take From This
Separate two questions that usually get merged: does the model know us, and does the model retrieve us. The first is about reputation and training data and there is very little you can do about it quarter to quarter. The second is about whether your pages rank for the question and are shaped like an answer when they do.
Almost every established brand we have benchmarked scores high on the first and low on the second, and almost every AI visibility program should be spending its effort on the second.
Benchmark by topic and by engine before you change anything. A single "AI visibility score" hides the pattern that tells you what to do.
Four engines at zero on the same topics is one problem, usually ranking or crawler access. One engine at zero while others cite you is a different problem, usually content shape or a blocked bot. You cannot tell them apart without per-engine data, and you cannot get per-engine data without tracking that runs the same prompts on a schedule.
Then build the answer scaffolding into whatever content process you already run, rather than standing up a parallel AI content stream.
An answer capsule, question-led headings, an FAQ that answers real questions, valid schema, and cluster links are cheap to add to a page you are already touching and expensive to retrofit later. If a system can generate and validate them for every page automatically, the cost approaches zero and the coverage approaches complete.
And hold your nerve for six weeks. Retrieval engines cite what they can find, and they can only find pages that have climbed the index. The three flat weeks were the hardest part of this program to sit through and the least surprising in hindsight.
If you want to know where your brand is cited, where it is not, and which of those gaps are fixable this quarter, we can run the same per-engine benchmark on your own prompts and walk you through it.
Bring this dispatch into a working session - one page in, scoping memo out.
Brief Foyer
