The entire GEO industry is built on a measurement system that doesn't exist. So we ran the test it's been avoiding. The results are awkward.

Agencies sell "AI visibility audits." SaaS tools promise to "track your AI rankings." Conference talks, white papers, LinkedIn threads, all full of confident claims about what works. Ask any of them a simple question: when someone asks an AI for a CRM recommendation, which source does it actually cite? The exact domain named in the answer. The URL that nets the click-through. Ask any three agencies the same question and you'll get three different methodologies, none of them publishable.

We asked three agencies and got three methodologies, none of them published. So we built the measurement and ran it.

Field Notes Why Is Nobody Measuring Who AI Actually Cites? GEO DataDab.com

What We Actually Ran

Two hundred decision-stage queries across five B2B categories: CRM, email marketing, project management, analytics, and customer support. The kind of question a real buyer asks. "HubSpot vs Salesforce for a 200-person team." "Best CRM for mid-market SaaS." "Alternatives to Amplitude." Comparison and buying-criteria queries, the ones that trigger vendor recommendations.

Every query went through one system first out of the gate: Perplexity. ChatGPT and Gemini citation streams aren't exposed the same way through plain APIs, so their overlap measurement lands in a future edition. We state that limitation plainly.

What came back was a surprise.

Every single query returned citations, every one of the two hundred. That's 1,896 citation URLs in total, an average of 9.5 per query. No empty answers, no "I can't browse the web" dead ends. Perplexity cites something in every response.

And the most-cited source in the whole study wasn't a category leader. It was Reddit.

Reddit appeared in 48% of all 200 queries. It held the single most-cited domain in every one of the five categories. Top source across the board.

The most-cited source is Reddit. Top cited domains across 200 Perplexity decision-stage queries reddit.com 96 g2.com 33 linkedin.com 24 zapier.com 17 featurebase.app 16 monday.com 16 Source: DataDab measure, 200 queries / 5 categories, Perplexity, 2026-08-16

The Brands Winning Have No Idea They're Winning

Here's the awkward part for the GEO industry. The sources Perplexity reaches for are missing from anyone's optimisation budget.

Across the study, the most-cited domains were Reddit (96 citations), G2 (33), LinkedIn (24), Zapier (17), and a string of comparison and review sites you've probably never heard of. Emailtooltester. Featurebase. Sequenzy. Folk. Niche properties with no "AI visibility strategy" and no vendor budget behind them. They're winning AI citations by accident, because they answer a buyer's question directly in plain HTML an LLM can parse.

The vendors spending six figures on "AI-optimised content" are chasing a system that cites Reddit threads they cannot edit and review roundups they did not commission.

This Is Not a Fluke of One Category

If Reddit dominated a single category, you could dismiss it as a quirk. It didn't. It was the top cited domain in all five.

Project management, 57.5% of queries cited Reddit. Analytics, 47.5%. Customer support, 47.5%. CRM, 45%. Email marketing, 42.5%. The pattern is consistent enough to call it behaviour.

It also lines up with how the system works: real-time retrieval, recency-weighted, leaning on forums and product discussion threads. The citation habits map to the architecture.

What a Measurement Gap Actually Costs

The immediate cost of not measuring is wasted spend. A SaaS company budgeting for "AI-optimised" posts while Reddit and review sites do the heavy lifting is allocating against a pipe dream. You would never optimise for Google without checking what Google ranked. The same discipline applies here.

The deeper cost is strategic blindness. Without a baseline you can't iterate, you can't test, you can't tell whether anything moves. The optimisation loop breaks at step one.

The strangest cost sits in plain sight: the brands actually winning AI citations today have no strategy at all. They're structurally positioned to get cited. Clear product pages, comparison content, documentation. Ask them what their GEO strategy is and they'll look confused. The people who'd benefit most from this data are the ones collecting the least of it.

BUILD THE SCORE CARD THE LEADERBOARD NOBODY'S BUILT

The Missing Leaderboard

What this space needs is exactly what SEO had in its early days: a public, regularly updated leaderboard. A plain page, free for anyone. Here are the twenty most-cited sources in Perplexity for CRM recommendations. Here are the twenty for email marketing. Here's how they moved from last quarter. Here's the raw data and the methodology, so anyone can replicate the measurement.

The data exists. AI responses are text. The citations are in the text. You can scrape them, count them, publish the results. Agencies have little incentive to do that. A public measurement lets clients verify the claims, which is uncomfortable for anyone selling "AI visibility" as a bespoke service.

Why This Changes Everything

Once you can measure, the conversation flips from "how do I get cited by AI?" to "who is actually getting cited, and why?" Those are different questions with different answers.

You start optimising for a system you can actually observe. You start comparing your citation share against competitors on a public scorecard. You start treating GEO as the measurement problem it always was.

If the first leaderboard's findings hold, the uncomfortable answer is that the current winners are Reddit threads and five-person review blogs. A B2B brand can skip trying to out-optimise Salesforce entirely. Win by becoming the first genuinely useful, AI-extractable source in a category full of marketing fluff. That's a strategy you can't buy, only build.

Get access to the raw dataset and methodology.

Learn more

The Methodology Agencies Prefer to Keep Vague

Here's what a credible leaderboard needs to actually work: an open method somebody else could replicate on a Saturday afternoon, free of any proprietary black-box scoring.

Fig. 1 — Measurement Pipeline 01 Query Set 200-500 prompts 02 Run Each System 48hr window 03 Record Citations full URL + domain 04 Score: citations per query, per domain, per system aggregated quarterly Leaderboard: top 20 per category per system DataDab Open Methodology v0.1

The core mechanic is straightforward. Pick a set of queries, run them through the system, record every source it cites, count. The complexity lives in the choices. For the first edition we measured one system, Perplexity, across 200 queries. ChatGPT and Gemini citation streams arrive in a later edition. Better to publish real single-system numbers than to pretend at a three-way comparison we didn't run.

First, the queries. Two hundred across five categories, all decision-stage. Informational prompts like "what is CRM?" produce citation patterns that look nothing like buying questions, so we left them out. We used the questions that actually trigger vendor recommendations: comparisons, "best X for Y", "alternatives to Z".

Second, the system. Perplexity, because its API returns structured citations we could read directly. Each query ran through its real search mode, so the citations reflect what an actual user sees.

Third, the recording. Every URL in every response was logged from the structured citation field, full path included. A citation of mixpanel.com/product reads differently from mixpanel.com/blog/what-is-product-analytics. The page matters as much as the host.

Fourth, the scoring. Citations per query, aggregated per domain, per category. That's the leaderboard. The methodology is boring on purpose, because boring methodology is reproducible, and reproducible methodology is the only kind that earns trust.

Reddit dominates every category. Share of queries citing Reddit, per category (higher = more Reddit citations) Project mgmt 57.5% Analytics 47.5% Support 47.5% CRM 45.0% Email 42.5% Source: DataDab measure, 40 queries per category, Perplexity, 2026-08-16

The Categories, Measured

Categories behave differently, and the differences are where the strategy lives.

CRM. Reddit showed up in 45% of queries. The surprise sat on the brand leaderboard: Nimble (15 citations) and Insightly (13) out-cited Salesforce (7). The tools everyone assumes own this space barely register in the cited-domains count. Being huge doesn't make you the reference.

Email marketing. Reddit in 42.5% of queries. The real story is a cluster of tiny comparison and automation sites most marketers skip: Sequenzy (16) and Campaign Monitor (16), then Braze (12), Mailchimp (10), with G2 and Emailtester.com in the mix. Mailchimp, the household name, isn't the citation winner its awareness would suggest.

Project management. The highest Reddit rate of the five at 57.5%. Atlassian's stable is strong as an ecosystem at atlassian.com (15), with Asana (14), Wrike (12), and Monday (11) close behind. Reddit threads and comparison guides carry the category more than any individual vendor's marketing.

Analytics. Reddit in 47.5%. Mixpanel leads organic domains at 19, with the open-source Matomo (13) ahead of Adobe (11) and Amplitude (10). A privacy-friendly open-source tool out-citing two enterprise giants is a genuinely useful data point for the PLG crowd.

Customer support. Reddit in 47.5%. Help Scout (15) leads, ahead of Featurebase (14) and Kayako (13). Zendesk, the category's default name, sits at 9. The best-regarded tool according to citations is helpfully smaller than the biggest.

Five categories, one clear pattern: the platforms and the genuinely useful niche players win, and the default enterprise names underperform. That's a finding a brand can act on.

Why the First Edition Will Be Wrong on Purpose

Every measurement system is wrong in its first iteration. The question is whether it's wrong in a useful way.

The first DataDab leaderboard is deliberately narrow. One system, 200 queries. It will miss citations that appear only in specific conversation threads, and the long tail of queries we didn't ask. It undercounts, on purpose.

That's fine. The first edition is transparent about its own limits: here's what we measured, here's how, here's what we found, here's what we missed. Precision can wait. The editor who publishes the numbers at all wins the argument.

The second edition adds ChatGPT and Gemini citation streams and settles the overlapping-source question properly. The third edition widens further. By the fourth, the methodology will have been challenged, refined, and forked by other researchers. That's how measurement becomes standard: it gets opened up and improved, and the first version has to be publishable enough to start that argument.

The leaderboard needs to be the first word somebody says with evidence behind it, and it has to start somewhere imperfect.

Want the raw dataset and methodology? The full tables, every prompt, and the open method live on the Perplexity leaderboard. Next edition adds ChatGPT and Gemini citation streams.

Learn more