AI answer engines (ChatGPT, Perplexity, Google AI Overviews) are capturing a growing share of informational searches, and no standard measurement tool natively tells you whether your content is cited there, or how much traffic arrives from it.
The problem breaks down into independent signals no single metric solves: technical visibility to crawlers, measurable referred traffic, real bot crawling, and cross-referencing both signals.
A 5-layer framework on this very site: technical foundation, referred traffic in GA4, bot crawling, cross-referencing both signals in BigQuery, and real citation testing in answer engines.
All 5 layers verified in production. The traffic cross-reference doesn’t find overlap yet, expected at this volume, and the citation test comes back at 0% across 80 real checks: the brand doesn’t show up in AI answer engines yet.
Context
Before selling GEO/AEO measurement to a client, it had to be applied to myself first. This site is the testbed: each layer of the framework gets built, verified in production with real evidence (console captures, HTTP headers, Search Console reports), and only then documented here. Nothing on this page describes something that isn’t actually deployed.
How it was actually solved
The 5 layers, one by one
Layer 1, verified: robots.txt explicitly allowing GPTBot, ClaudeBot, PerplexityBot and similar, TechArticle and FAQPage structured data on every page, content written to be citable. Layer 2, verified: a Custom Channel Group in GA4 that isolates identifiable referred traffic from AI platforms using Google’s native classification. Layer 3, verified: Cloudflare’s AI Crawl Control, natively active on the domain, logs every AI bot visit (Amazonbot, Googlebot, GPTBot, ClaudeBot, PerplexityBot and dozens more) without blocking any of them. Layer 4, verified: a script queries Cloudflare’s GraphQL Analytics API daily (the AI Crawl Control panel has no export on the free plan), loads it into BigQuery next to GA4’s native export, and a JOIN by page and date cross-references "crawls me" with "sends me traffic". Layer 5, verified: a script runs the same 20 prompts against Gemini with Search Grounding on, the same signal a real user sees, and logs whether the brand shows up and which domains the model actually cites.
Why crawling, referred traffic, and citation are three different things
An AI bot crawling your page doesn’t mean it cites you in an answer, and being cited doesn’t mean the user clicks through with an identifiable referrer. Many AI apps (ChatGPT’s mobile app, a good chunk of Perplexity’s traffic) send no Referer at all, so that traffic falls into "(direct)", indistinguishable from real direct traffic. The framework measures all three signals separately precisely because collapsing either of the first two into the third hides most of the real activity.
Why Layer 3 is logging, not blocking
Blocking AI crawlers, something this same site did deliberately, at a different point, against abusive scraping bots, defeats the purpose when the bot in question is the one deciding whether to cite you. GPTBot and ClaudeBot need to be able to read the content to cite it. Verified in the Cloudflare dashboard: in the first 24 hours of review, 29 AI crawler requests, all allowed, none blocked.
What the first real cross-reference shows, and what it doesn’t
With 3 days of data, the JOIN doesn’t yet find a single page where a session with an identifiable AI source lines up with an AI bot crawling it the same day. At this traffic volume and such a short window, that’s expected, not a conclusion about whether the content performs in answer engines. The script that loads the Cloudflare data doesn’t run on a cron yet either: it’s run by hand until it’s worth automating.
Why 0% citation across 80 checks isn’t a measurement failure
80 checks, 60 manual in August and 20 automated since, and none cite the brand. The domains Gemini cites instead (Malt, agencies with years of history, job boards) share something a new portfolio site doesn’t have yet: external authority, backlinks, mentions on other sites. Automating the same check on ChatGPT, Claude and Perplexity wouldn’t change that conclusion, so it was shelved for now. This isn’t a measurement problem, it’s an authority problem measurement alone doesn’t solve.
GEO/AEO is too new a field for a single standard metric to exist the way it does in classic SEO. Any serious framework has to break the problem into separately measurable signals, and be honest about which of those signals are actually solved and which aren’t, even when the first real result is an empty cross-reference or a 0% citation rate.
FAQ
Is this framework finished?
All 5 layers are built and verified with real data. What’s missing is signal: only a few days of traffic for Layer 4’s cross-reference, and 0% real citation across Layer 5’s 80 checks, consistent with a new site’s lack of external authority.
Can GEO/AEO be measured without access to Cloudflare or BigQuery?
Partially: a Custom Channel Group in GA4 (Layer 2) already isolates some referred traffic using tools almost any account already has. Seeing which AI bots crawl you (Layer 3) and cross-referencing that with real traffic (Layer 4) does require access to CDN logs and a data warehouse like BigQuery.