Google cited cleanlist.ai in 30 of 58 live AI answers on B2B data buyer queries on August 10, 2026, and wrote the word Cleanlist in exactly one of them. In the 22 Google AI Mode answers, cleanlist.ai was the single most-cited source in the set at 14 of 22, and Cleanlist was named zero times. Twenty-nine of those thirty citations are ghost citations: Google read a Cleanlist page and then recommended ZoomInfo, Apollo, Clay, Cognism or HubSpot in the text the buyer actually reads. The full keyword-level dataset is published below under CC BY 4.0.
Last updated: August 10, 2026. All 58 answers were collected live on August 10, 2026.
Cleanlist ran this study and Cleanlist is the subject of it
Cleanlist is a B2B data enrichment company. We picked the keywords, ran the collection, wrote the counting code and are publishing the result. The keywords are drawn mostly from our own AI Overview citation gap and ranked by volume multiplied by CPC, so this is a sample of one vendor's footprint rather than a random sample of the category.
Two things reduce how much you have to take on trust. The 37-keyword dataset is published in full, so every count here can be recomputed without us. And the counting rules are printed on the page, including the one that cut our own naming figure from 6 to 0 before publication.
What is a ghost citation?
A ghost citation, the term Cleanlist uses throughout this August 10, 2026 report, is an AI answer that reads your page and recommends somebody else. The engine pulls your URL into its source set, prints your name on the citation card, then writes a visible answer naming a different vendor. Cleanlist measured 58 Google AI answers on August 10, 2026 and found 30 citations of cleanlist.ai against 1 mention of the Cleanlist brand in answer text, a ratio of 30 to 1. The distinction matters commercially because a buyer scans the answer and rarely opens the sources. Being retrieved is a page-level outcome. Being named is a brand-level outcome, and the two are produced by different work.
How often does Google cite a source without naming the company behind it?
On this sample, almost always. Across 36 live Google AI Overviews collected on August 10, 2026, cleanlist.ai was cited in 16 (44%) and Cleanlist was named in the visible answer in 1 (3%). Across 22 Google AI Mode answers on the same day, cleanlist.ai was cited in 14 (64%), the most-cited source label in the set, and Cleanlist was named in 0. Combined, that is 30 citations and 1 name across 58 answers, a 97% ghost rate for this vendor on these queries.
| Surface | Answers | cleanlist.ai cited | Cleanlist named |
|---|---|---|---|
| Google AI Overview | 36 | 16 (44%) | 1 (3%) |
| Google AI Mode | 22 | 14 (64%) | 0 (0%) |
| Both | 58 | 30 | 1 |
Which domains does Google actually cite in AI Overviews for B2B data queries?
YouTube. Across the 36 AI Overviews, YouTube was cited on 19 keywords, ahead of Cleanlist on 16, ZoomInfo's blog on 13 and Reddit on 12. Those answers drew on 144 distinct source labels across 305 citation slots, so the tail is long, but a video platform sits above every B2B data vendor in the category Cleanlist competes in. Cleanlist has no YouTube presence as of August 10, 2026, which makes this the largest structural gap in our own profile and the clearest actionable finding here.
| Source | AI Overviews citing it (of 36) |
|---|---|
| YouTube | 19 |
| Cleanlist | 16 |
| ZoomInfo Blog | 13 |
| 12 | |
| Cognism | 9 |
| Autobound.ai | 8 |
| Apollo.io | 7 |
| Salesgenie | 6 |
| HubSpot | 6 |
Which sources does Google AI Mode cite for the same queries?
A different set, and a smaller-company set. In the 22 AI Mode answers of August 10, 2026, Cleanlist was cited on 14 keywords, SyncGTM on 13, YouTube on 12, ZoomInfo's blog on 9 and Reddit on 8, across 148 distinct source labels and 299 citation slots. Several domains matching or out-citing the incumbents are small AI-native GTM sites: SyncGTM, Autobound.ai, Amplemarket, Landbase, Salesmotion and MarketBetter each appear on five or more of the 22 keywords.
| Source | AI Mode answers citing it (of 22) |
|---|---|
| Cleanlist | 14 |
| SyncGTM | 13 |
| YouTube | 12 |
| ZoomInfo Blog | 9 |
| 8 | |
| Clay | 7 |
| Autobound.ai | 7 |
| Cognism | 6 |
| Demandbase | 6 |
| Amplemarket | 6 |
Which brands does Google name in the visible answer?
The incumbents, by a wide margin. Counting brand names in answer text after every source URL and citation block was stripped out, ZoomInfo was named in 22 of the 36 AI Overviews and Apollo in 20, then HubSpot 18, Clay 17, Salesforce 15, Cognism 11 and Lusha 10. In AI Mode it is ZoomInfo 14 and Apollo 14 of 22, then Salesforce 13, HubSpot 13, and Clay and Cognism at 11 each. Cleanlist was named once across all 58 answers. These are mention counts: a brand named as the thing to avoid counts the same as one named as the pick. Three rows were corrected before publication, all from the same cause, a brand name that is also an ordinary English word. Instantly matches the adverb, and every one of its hits across the 58 answers was the adverb, which takes Instantly from 6 and 3 to 0 and 0. The same effect on "seamless" took Seamless.AI from 3 and 3 to 2 and 2. And the AI Mode answer for clay pricing is about ceramics rather than software, so that hit comes out of Clay's AI Mode column, taking it from 12 to 11.
| Brand | Named in AI Overviews (of 36) | Named in AI Mode (of 22) |
|---|---|---|
| ZoomInfo | 22 | 14 |
| Apollo | 20 | 14 |
| HubSpot | 18 | 13 |
| Clay | 17 | 11 |
| Salesforce | 15 | 13 |
| Cognism | 11 | 11 |
| Lusha | 10 | 7 |
| Clearbit | 5 | 7 |
| UpLead | 6 | 2 |
| Seamless.AI | 2 | 2 |
| Instantly | 0 | 0 |
| Cleanlist | 1 | 0 |
What does a ghost citation look like on a single keyword?
It looks like this. On data enrichment platform, Cleanlist ranks 2nd organically, the AI Overview cites cleanlist.ai, and the answer names ZoomInfo, Apollo, Clay and Cognism. On lead enrichment api, where Cleanlist ships a public enrichment API, the AI Mode answer cites cleanlist.ai and names ZoomInfo, Apollo, Clearbit, HubSpot, Salesforce, People Data Labs and Autobound. Each of these keywords was collected on August 10, 2026 and is a row in the published CSV.
| Keyword (AI Mode) | Our page cited | Brands named in the answer |
|---|---|---|
| contact data enrichment | yes | ZoomInfo, Apollo, Clay, Cognism, Lusha, Clearbit, HubSpot |
| data enrichment software | yes | ZoomInfo, Apollo, Clay, Cognism, Lusha, Clearbit, HubSpot |
| email enrichment | yes | ZoomInfo, Apollo, Clay, Cognism, Hunter.io |
| enrich contacts | yes | ZoomInfo, Apollo, Clay, Cognism, Lusha, Clearbit, HubSpot |
| lead enrichment api | yes | ZoomInfo, Apollo, Clearbit, HubSpot, Salesforce, People Data Labs, Autobound |
What did the one answer that named Cleanlist actually say?
It said this, in the AI Overview for b2b data api on August 10, 2026, inside a provider comparison table: Cleanlist, "High-accuracy data enrichment", "Waterfall querying across 15+ sources for ~98% fill rate". Two things are worth stating about that. The 98% figure traces to the Cleanlist 500-Lead Enrichment Benchmark 2026, which measured 98% verified email and 85% direct dial across 500 stratified B2B leads. The "15+ sources" figure is ours and it is unsourced. Cleanlist's waterfall runs 15+ providers. Google repeated a number it found on our own pages, which is the whole reason a vendor's published numbers have to be correct.
Which Cleanlist pages is Google reading?
Very few of them, which is the useful part. Across the 36 AI Overviews, Google cited 8 distinct cleanlist.ai URLs, filling 30 citation slots. One page, the best data enrichment tools guide published March 31, 2026, accounts for 6 of the 15 keywords with a resolvable citation URL and 13 of the 30 slots on its own. The other seven URLs contribute between 2 and 5 slots each. This matches the fan-out pattern: an engine that decomposes a query into sub-questions tends to return to one page it already trusts rather than spreading citations across a site.
| Cited URL | Keywords | Citation slots |
|---|---|---|
| /blog/2026-03-31-best-data-enrichment-tools-2026 | 6 | 13 |
| /blog/2026-04-20-best-b2b-contact-database-software | 3 | 5 |
| /blog/2026-03-19-rocketreach-pricing-guide | 1 | 2 |
| /blog/2026-03-19-zoominfo-pricing-guide | 1 | 2 |
| /alternatives/zoominfo | 1 | 2 |
| /blog/2026-03-19-seamless-ai-pricing-guide | 1 | 2 |
| /blog/2026-03-19-salesforce-data-enrichment-guide | 1 | 2 |
| /blog/2026-05-12-best-database-cleaning-software | 1 | 2 |
Does ranking organically make a citation more likely?
It helps, and it is not required. Of the 16 AI Overviews that cited cleanlist.ai on August 10, 2026, 8 sit on keywords where Cleanlist ranks in the top 4 organically, 4 where it ranks between 8 and 18, and 4 where Cleanlist has no top-20 position at all: zoominfo pricing, zoominfo competitors, data enrichment services and salesforce data enrichment. The reverse happens too. On lead list building Cleanlist ranks 11th and the AI Overview cites Apollo, Hunter.io, Instantly, LinkedIn and YouTube instead. Retrieval and ranking overlap without being the same system.
How big is the citation gap across a whole site?
Larger than the sample suggests. A Semrush US organic positions export dated August 9, 2026 shows 4,261 unique keywords where cleanlist.ai ranks. 3,146 of them (73.8%) carry an AI Overview, and cleanlist.ai is cited in 400 of those (12.7%). That leaves 2,691 keywords where Cleanlist ranks, an AI Overview exists, and Cleanlist is absent from it: 917,550 monthly searches and $5.53M per month in equivalent paid-traffic value at published CPCs. 586 of those gap keywords already carry a top-10 organic position, covering 132,740 monthly searches, which is the cheapest slice to work on first.
Which keywords are worth the most in that gap?
Ranked by volume multiplied by CPC rather than by volume, the gap reads as a buyer list. The top of it on August 9, 2026 is rocketreach pricing (27,100 searches, $42.36 CPC), zoominfo competitors (33,100, $32.27), zoominfo pricing (27,100, $18.82), then how to verify email address (110,000, $1.38) and clay pricing (1,900, $62.49). Below that sit low-volume, very high-CPC buyer terms: what is lead enrichment at 260 searches and $212.00 CPC, contact data enrichment at 480 and $96.17, hubspot data enrichment at 480 and $82.30. Cleanlist ranks on all of them, and the August 9 export records no AI Overview citation for any of them. The August 10 live pull disagrees on four: rocketreach pricing, zoominfo competitors, zoominfo pricing and contact data enrichment were all cited that day. A gap measured from a weekly export is a week-old view of a surface that moves daily.
Why did Cleanlist recount its own numbers before publishing this?
Because the first pass got them wrong in our favour, in a familiar way. An intermediate summary of this pull reported 13 AI Mode citations and 6 AI Mode brand mentions for Cleanlist. Recounting from the archived API responses gives 14 citations and 0 mentions. All 6 of those mentions were our own sentences read back to us: Google attaches a preview snippet to each citation object, lifted from the cited page, and every one of the 6 hits sat in that snippet rather than in the answer Google wrote. Cleanlist published a correction for the same family of bug on August 7, 2026, three days earlier, where the string turned up inside cited link targets instead. Finding it a second time in a new script is the reason the counting rules are printed on this page.
The two counting fixes applied before publication
Fix 1, naming. Brand names are counted only in the visible answer body, after every citation block is removed and every URL is stripped from what remains. Both steps are load-bearing. Google ships a preview snippet inside each citation object, copied verbatim from the cited page, so the card for cleanlist.ai/blog/2026-05-22-best-zoominfo-alternatives-2026 carries our own sentence "ZoomInfo's main competitors in 2026 are Cleanlist, Apollo.io, Cognism, Lusha, Clay, Hunter.io, LeadIQ, and 6sense" into the response payload. All 6 of Cleanlist's AI Mode naming hits were that snippet text rather than anything Google wrote. Dropping the citation objects takes the count from 6 to 0; the URL strip catches the separate case of a brand name sitting inside a cited link target. The bias runs one way. It inflates whichever brand publishes the pages the engine cites, which is disproportionately the brand running the measurement.
Fix 2, citations. Google sometimes returns AI answer references as redirect URLs whose domain is literally google.com. A counter that reads only the domain field silently drops those keywords. On b2b data api and rocketreach pricing the citation is visible only in the source label Google prints on the card. Counting the label as well as the domain takes the AI Overview total from 15 to 16 and the AI Mode total from 13 to 14.
The naming fix costs us six mentions and the citation fix gains us two. Both are applied because they are correct, whichever way they point.
Why would an engine cite a page and name a competitor?
Here is the mechanism we believe drives it, stated as a hypothesis this data is consistent with rather than a finding it proves. A passage explaining a category without naming a vendor is extractable and unattributable. "Waterfall enrichment queries providers in sequence until a verified result returns" is a perfect sentence to lift, and it carries no brand, so the engine attaches whichever vendor names it already associates with the category. Written as "Cleanlist's waterfall queries 15+ providers and charges only when a verified result returns", the same claim cannot be lifted without the name. We have not run the controlled test that would confirm this.
What are the limits of this study?
Six, stated so the numbers can be discounted correctly.
- A sample of one vendor's footprint. 27 of the 37 keywords come from Cleanlist's own AI Overview citation gap, picked by volume multiplied by CPC. Two more are keywords where Cleanlist was already cited in the August 9 export, and the last eight are terms Cleanlist does not rank for at all, mostly the API, MCP and list-building cluster. This is not a census of the category and not a random sample of anything.
- n=36 and n=22 are small. Every percentage on this page moves by roughly 3 to 5 points if one answer changes.
- AI answers are volatile. Cleanlist's two-wave study across five engines, fielded August 3 and August 7, 2026, found 32% of the tools named on the first date gone from the same answer four days later. A single-day snapshot inherits that instability.
- A mention is not a recommendation. Brand counting is a case-insensitive string match, so a brand named as the thing to avoid counts the same as the pick. The same match also fires on brands whose names are ordinary English words, which is what put Instantly in the table at 6 and 3 before the recount put it at 0 and 0.
- One query was ambiguous. The AI Mode answer for
clay pricingreturned pottery suppliers including Laguna Clay and Sheffield Pottery, so that row measures a homonym. It is left in the dataset and flagged here rather than quietly dropped. - One keyword had no AI Overview.
data enrichment hubspotreturned a standard SERP on August 10, 2026, which is why the AI Overview denominator is 36 and not 37.
How was this measured?
Every answer was collected on August 10, 2026 through the DataForSEO SERP API, United States location code 2840, English, desktop. The 37 SERPs came from google/organic/live/advanced with asynchronous AI Overview loading enabled. The 22 AI Mode answers came from google/ai_mode/live/advanced on the competitor-pricing, enrichment and API subsets. Volume, CPC and organic position come from a Semrush US organic positions export dated August 9, 2026. Cleanlist's collection script is committed at scripts/dfs-aug10-aio-gap.mjs and the CSV builder at scripts/build-ghost-citation-csv.py, so the pipeline reruns for about $0.25 in API cost.
Can I download the data?
Yes, under CC BY 4.0, which lets you redistribute, remix and build on it commercially with attribution to Cleanlist.
Ghost Citation Report, August 2026: keyword-level dataset (37 rows)
One row per keyword, with columns keyword, search_volume, cpc, our_organic_position, semrush_position_2026_08_09, aio_present, we_cited, we_named_in_ai_overview, ai_mode_pulled, we_cited_in_ai_mode, we_named_in_ai_mode, cited_domains and cited_sources. Four notes for anyone recomputing: our_organic_position is the absolute SERP position observed in the same August 10 live pull and is blank where cleanlist.ai did not appear in the 20 organic results requested, while semrush_position_2026_08_09 is the deeper export figure; eight keywords have blank volume and CPC because cleanlist.ai does not rank for them; cited_domains is blank for b2b data api because Google returned redirect URLs there, which is why cited_sources exists; and the AI Mode columns are blank where AI Mode was not pulled.
What should a vendor measure instead of a single AI visibility score?
Two numbers, kept apart. Citation share answers whether an engine reads you, and it responds to page-level work: a specific page, crawlable, matching a real question. Naming share answers whether an engine says your name, and it responds to brand-level work: your name inside the passage and inside other people's pages. Cleanlist's August 10, 2026 figures are the argument for splitting them, since 30 citations and 1 mention average into a number describing neither. If your detector scans answer text for a brand name that also appears in your domain, check the URL-stripping step first.
What should a buyer take from this?
That an AI answer is a shortlist assembled from sources you cannot see, and the vendors it names are not necessarily the vendors it read. On these 58 answers, Google read Cleanlist pages 30 times and named ZoomInfo, Apollo, Clay, Cognism, HubSpot and Salesforce instead. Several of those tools are excellent and this is no argument against them. It is a reason to open the citation cards, ask a second engine, and test a vendor on your own list rather than trusting one paragraph. Cleanlist's free plan is 30 credits a month with no card and paid plans start at $79 a month, enough to check a real list.
Test it on your own list
Upload a CSV and run it through the Cleanlist waterfall in bulk. Free plan is 30 credits a month, no credit card. Search costs 0 credits, and enrichment is pay-for-results, so you are not charged when no data is found.
FAQ
What is a ghost citation in AI search?
A ghost citation is an AI answer that cites your page as a source and names a different company in its visible text. Cleanlist measured 58 live Google AI answers on B2B data buyer queries on August 10, 2026: cleanlist.ai was cited in 30 of them and the word Cleanlist appeared in the answer body of 1. The other 29 citations are ghost citations. The term matters because most AI visibility tooling reports a single blended score, and a vendor with high citation share and near-zero naming share has a very different problem from a vendor with the reverse. Both are measurable separately from the same raw answer.
How often do Google AI Overviews name the sites they cite?
On this sample, rarely for the publisher itself. Across 36 live AI Overviews collected on August 10, 2026, cleanlist.ai was cited on 16 keywords and Cleanlist was named in the answer body on 1. In the same 36 answers ZoomInfo was named 22 times and Apollo 20 times, while ZoomInfo's blog was cited 13 times and apollo.io 7 times, so for those brands the naming count runs ahead of the citation count. The pattern is that engines name the brands their sources talk about, and cite the pages that do the talking. This is a 36-answer sample on one vendor's keyword set and should not be read as a general law.
Does Cleanlist's own research make Cleanlist look good?
No, and that is the point of publishing it. The headline finding is that Google reads Cleanlist pages constantly and almost never says the name: 30 citations and 1 mention across 58 answers on August 10, 2026. The report also documents that our first pass at these numbers overstated Cleanlist's naming count as 6 rather than 0, through a second variant of the self-citation bug Cleanlist published a correction for on August 7, 2026. The dataset is open under CC BY 4.0 and the collection script is committed, so the weak numbers can be verified as easily as the strong ones.
Can I reproduce this study?
Yes. The Cleanlist keyword-level dataset is published under CC BY 4.0 at cleanlist.ai/data/ghost-citation-report-2026-08.csv and the collection script is committed at scripts/dfs-aug10-aio-gap.mjs. Everything was collected through the DataForSEO SERP API on August 10, 2026, United States location code 2840, English, for roughly $0.25 in API cost. Swap in your own keywords and apply the same two rules: strip URLs and citation blocks before matching brand names, and read both the domain and the source label when counting citations. If your result contradicts ours, publish it.
Run this on your own list
Upload a CSV, enrich every row across 15+ providers, and export the result back to your CRM. First 30 rows free.
Enrich 30 rows free30 credits free every month · No credit card
