If someone asks ChatGPT or Perplexity which tool to buy in your category, and the answer cites you, that visitor lands on your site. The question almost nobody can answer today is: how much of that traffic is actually being measured?
The problem isn’t visibility, it’s measurement
There are agencies selling “GEO optimization”: tweaking content so language models cite it more. It’s a real discipline. But optimizing without measuring is like driving with your eyes closed: you don’t know if what you’re doing works, or how much business it generates.
The 4 layers you need
1. Visibility: does your brand show up when someone asks about your category? A prompt set runs recurringly against the major models, tracking position and which sources get cited.
2. Referred traffic: sessions arriving from LLM interfaces need explicit capture in GA4. Without configuration, GA4 doesn’t distinguish a ChatGPT visit from a direct one:
// Example Channel Group regex for GA4
^(chatgpt\.com|chat\.openai\.com|perplexity\.ai|
gemini\.google\.com|claude\.ai)$
3. Crawlers: not all bots are equal. A training bot (GPTBot, ClaudeBot, CCBot) doesn’t correlate with getting cited — it just feeds a model that will be trained months later. A real-time retrieval bot (OAI-SearchBot, PerplexityBot, ClaudeBot in search mode) does matter, because it’s the one querying your page at the exact moment someone asks something. Telling them apart by user-agent in your server logs (or in Cloudflare’s bot categories, if you’re already managing AI bot access there) is the only way to know whether you’re being cited live or just archived for later.
4. Attribution: the layer almost nobody delivers. Joining the three above with your real BigQuery pipeline, to answer which prompts cite you and how much revenue those sessions generate.
What layer 1 actually looks like in practice (a real data point, not a made-up example)
To not stay purely theoretical: I run this exact layer-1 pilot on my own site, with 20 real prompts of the kind a potential client would use (“freelance GA4 and GTM consultant in Madrid”, “who can audit my Google Tag Manager container”…), against Gemini, saving every structured response to BigQuery. Three rounds across three different days: 0 out of 60 responses mention me. Zero.
That’s not a failure of the method — it’s exactly the kind of result you need to be able to measure. The interesting part isn’t the zero: it’s which domains do get cited for those same questions (niche GEO tools like Peec AI or Otterly.AI, platforms like malt.es or stape.io). That’s already actionable information — it tells you what the model currently treats as authoritative, and it’s the real starting point for any content strategy, not a vanity number.
Why layer 4 is the one that actually matters
The first three layers say “you’re being mentioned.” Only the fourth says “this is making you money”, and that’s the difference between a vanity metric and a real case for investing in this.
An example of the kind of work behind this layer: reconciling a Google Ads vs. GA4 discrepancy by dropping to the raw event table in BigQuery, until the gap breaks down into exact, quantified causes. It’s documented as a real case: the same technical skill, applied here to a different question.
Want to know if your brand is already being cited, and whether that traffic is being measured? I can audit it.