A client asked me last month whether their “AI SEO strategy” was working. I asked what they meant by AI SEO. They weren’t sure. That mix-up happens constantly, and it costs marketing teams real time and real budget. “AI SEO” is actually two separate jobs wearing the same name tag.
Here’s why that distinction matters right now, what each job actually requires, and how to measure whether either one is working.
Why It Matters
I’ve watched dozens of clients pour effort into “AI visibility” without ever defining what they were trying to move. Some meant using AI tools to speed up keyword research and drafting. Others meant getting cited inside AI Overviews, Copilot, or ChatGPT search. Those are not the same project, and lumping them together blurs accountability for both.
The practical cost shows up fast. Teams report a single “AI visibility score” to leadership, can’t explain what actually changed, then burn a quarter chasing tactics that don’t work the way people assume — llms.txt files, invented schema, chunked paragraphs. If you’re a founder or marketer deciding where to spend the next sprint, you need to know which job you’re actually doing. That choice determines what evidence you should even be collecting.
The Two Jobs, Defined
The first job uses AI inside your existing SEO workflow. Think clustering queries, extracting entities from a known set of pages, flagging inconsistent terminology, drafting internal-link candidates, or generating test cases for a technical audit. The output here is a reviewed work product, and you judge it on accuracy, time saved, and how much a human had to fix afterward.
The second job makes a site technically eligible, genuinely useful, and — this is the part people skip — measurable across AI-mediated discovery surfaces like Google AI Overviews, AI Mode, Microsoft Copilot, and ChatGPT search. That’s closer to what people mean by AEO or GEO. Google’s July 2026 guidance on optimizing for generative AI features frames its AI surfaces as sitting on top of core Search ranking and quality systems, built around retrieval-augmented generation and what it calls query fan-out. A page still needs to be indexed and eligible to appear with a snippet in the first place. In other words: eligibility gets you a shot, not a guarantee.
I’ll be straight with you: this is the part that trips up even experienced teams. They nail the first job — AI speeds up their drafting pipeline — and assume that automatically buys them the second. It doesn’t. Different systems fail in different places.
The Foundation Hasn’t Moved
Here’s the reality nobody wants to hear because it’s not exciting: none of this replaces basic technical SEO. A tidy summary box at the top won’t rescue a page that’s blocked by robots rules, sitting behind a login, carrying a conflicting canonical tag, or excluded from the relevant index. Crawlability, index eligibility, internal linking, and a clear site structure remain the prerequisite, not a nice-to-have layered on afterward.
The content bar has actually gotten sharper, not softer. Google now pushes publishers toward what it calls “non-commodity” work: a distinct point of view, first-hand experience, or original evidence, rather than a longer rehash of what’s already out there. I’ve read plenty of genuinely comprehensive articles that still add nothing new. Comprehensive and valuable aren’t the same thing, and AI search seems to be widening that gap rather than closing it.
Other platforms run their own rulebooks. OpenAI documents separate crawlers for search discovery, potential model training, and user-requested page fetches. Allowing one bot tells you nothing about your standing with the others. I’ve watched teams assume one robots.txt tweak covered every surface at once — it doesn’t.
Why You Can’t Collapse These Into One Metric
This next part matters most for anyone actually trying to measure progress. AI search doesn’t move through a single funnel; it moves through a chain of separate states. Skip straight to “did we get cited” without checking the earlier steps, and you’ll mislead yourself.
Start with access: can the crawler even reach the page? Next comes index eligibility, then answer appearance, then a visible citation, then a referral click, then an actual conversion. Each stage needs its own evidence, and none of them proves the next stage happened. A page can sit in the index for months and never get retrieved for the query that matters to your business. It can get retrieved and never show up as a visible citation. It can get cited and never earn a click.
Microsoft’s own AI Performance documentation is blunt about this limit: its citation data shows visible source use across supported AI experiences, and it explicitly does not measure rankings, authority, or overall performance. Its grounding queries are grouped retrieval phrases too, not the literal prompts a person typed. OpenAI notes that ChatGPT referral links carry a UTM parameter that can help identify some visits, but that parameter still won’t surface every answer that referenced your page without sending a click.
What surprised me most, going through all this, is how many teams still report one blended “AI visibility” number to stakeholders. It feels efficient. It actually hides exactly where the funnel is breaking.
Building an Actual AI SEO Process
If you’re doing this for real, here’s roughly the sequence I’d walk a client through:
Define the reader’s decision, not the keyword. Get specific about what someone needs to decide after reading the page, and what facts or comparisons they need to make that decision.
Pick your surfaces deliberately. A test in Google AI Overviews on desktop in the US tells you nothing reliable about ChatGPT search on mobile in another market. Record country, language, device, and account state every time.
Confirm technical eligibility first. Check the status code, canonical, index directive, rendered text, sitemap inclusion, and internal links using the platform’s own webmaster tools before you touch content.
Add something that isn’t already out there. Build a worked example, an original test, a real dataset, or a transparent breakdown of what you actually observed versus what’s only documented. Teams skip this step because it’s the slow one.
Freeze a baseline before you touch anything. Save the URLs, queries, dates, screenshots, and known limitations before the edit, not after.
Change one variable at a time. Then measure index state, traditional rankings, citation reports, referral sessions, and conversions as separate line items. Resist the urge to merge them into a tidy composite score.
None of this guarantees a citation or a ranking bump. What it gives you is an honest record: what changed, what you actually know, and what you’re only hoping.
What to Skip Entirely

I’ve made my share of expensive mistakes chasing tool hype, so let me save you a few. Skip the llms.txt file if your goal is influencing Google Search rankings — Google says directly that it doesn’t use these files for ranking or visibility in its AI features. Skip chopping pages into artificially small “AI-friendly chunks,” too; no fixed paragraph length unlocks retrieval, and Google says no special chunking is required. Leave invented schema markup alone as well, since no dedicated schema type exists for AI Overviews. And stop calling a citation a conversion. A visible reference, a referral visit, and a completed action on your site are three different things, and they deserve three different rows in your spreadsheet.
Industry Impact
For developers and technical teams, this raises the stakes on the eligibility layer. Clean crawl access, correct canonicals, and working structured data now do double duty for both traditional rankings and AI surfaces, so technical debt costs more than it used to. Content and marketing teams should treat original evidence and first-hand testing as the real differentiator, not a nice bonus. Businesses reporting this to leadership need to resist compressing five different outcomes into one visibility score, even though that number fits a slide deck much more neatly.
Limitations and Open Questions
Independent, cross-platform benchmarking of AI citation behavior is still thin. A single test only produces a sample, not a guarantee of stable coverage, and results shift by locale, device, account state, and even the exact query path someone took to get there. Nobody, including the platforms themselves, claims a page can hold a fixed, universal “first place” inside a conversational answer.
What Comes Next
Watch how Google’s generative Search requirements and reporting evolve. Watch whether Microsoft expands what its AI Performance data actually discloses, and whether OpenAI adds clearer signals for what gets cited versus what merely gets crawled. Until then, teams that separate these two jobs, and keep separate evidence for each, will have a far clearer picture of what’s working than the ones chasing a single blended score.
The honest answer, as usual, depends entirely on what you’re actually trying to measure. Get specific about that first, and the rest of the process gets a lot less confusing.

