The Shortlist Census: What Google Recommends Across 240 B2B Data Buying Queries

240 buyer keywords and 60 AI Mode questions, Google US desktop, September 7, 2026. All 232 AI Overviews were readable. Google names a median of 3 vendors, and Cleanlist is on 6 of 232.

Victor Paraschiv

Victor Paraschiv

Co-Founder & CMO

31 min read

Google wrote a vendor shortlist above the results on 232 of 240 B2B data buying queries collected on September 7, 2026, which is 96.7%, and all 232 of those answers returned a readable body. They name 633 vendors in total, a median of 3 per answer. ZoomInfo is on 121 of them. Cleanlist is on 6, which is 2.6%, and on 1 of 60 Google AI Mode answers. In the same 232 answers, cleanlist.ai is the 7th most-read domain, cited in 32. We ran the identical query set a second time 25 minutes later and our own naming count came back 11 instead of 6, so read every Cleanlist naming figure on this page as the bottom of a 6-to-11 range and read the volatility section before quoting any of them.

Last updated: September 7, 2026. Every SERP and every AI Mode answer in this study was collected on that date.

Method

Collection and parameters. DataForSEO SERP API, Google Organic Live Advanced endpoint, one live pull per keyword, September 7, 2026, with the asynchronous AI Overview flag (load_async_ai_overview: true) set. The 60 buying questions were pulled the same day from a Google AI Mode endpoint. Both used google.com, location_code 2840 (United States), language English, device desktop, depth 50 for the organic block, no signed-in profile and no personalisation.

Sample frame. 240 buyer keywords in the B2B data, enrichment, prospecting and list-hygiene category, chosen by Cleanlist from a partner-filtered seed list, plus 60 buying questions phrased the way buyers type them.

Attempted, resolved, errors, bases. 240 keywords attempted, 240 resolved, 0 errors; 60 AI Mode questions attempted, 60 answered, 0 errors. 232 of the 240 keywords carried an AI Overview and all 232 returned a readable body, so every AI Overview rate divides by 232 and every AI Mode rate by 60. Where the attempted base of 240 gives a different figure, both are printed.

Which pull. The 240 keywords were collected twice, 25 minutes apart. Everything on this page except the volatility section is computed from the second pull, which is the pull published in sections 1 to 3 of the dataset. Organic positions come from the rank_group field, the position of a result inside the organic block, and never from rank_absolute, which counts every block on the page and understates a position on any SERP carrying an AI Overview.

Citation counting. Once per hostname per answer, however many times it appears in the source list. Hostnames are lowercased and a leading www. is stripped. Subdomains stay distinct in the domain leaderboard (pipeline.zoominfo.com and zoominfo.com are separate rows) and roll up to the registrable domain in the per-vendor tables, by the rule d == vd or d ends with . plus vd.

Naming counting. Once per vendor per answer. Every URL, markdown link target and bare domain is deleted from the answer text before any brand match runs, because a citation is not a naming. Brands are then matched on word boundaries against the lowercased remainder. The detector uses a fixed published vocabulary of 42 names, 32 of which appear here, so a vendor outside it scores zero by construction. Six ambiguous names (Clay, Instantly, Outreach, Apollo, Seamless.AI, Surfe) score only when a 440-character window around the match, 220 characters either side, carries a product word from a published list. That filter can undercount.

Keyword and question classes. Both class tables below are reproducible from column one of the dataset. A keyword takes the first class it matches, in this order: \bvs\b|\bversus\b gives comparison (vs); \balternatives?\b|\bcompetitors?\b gives alternatives / competitors; \bpricing\b|\bprice\b|\bcosts?\b|\bcheapest\b|how much gives pricing; a string starting how gives how to; \bbest\b|\btop\b gives best X; \btools?\b|\bsoftware\b|\bplatforms?\b|\bproviders?\b|\bvendors?\b|\bcompanies\b|\bservices?\b|\bapi\b|\bdatabase\b gives category / tools; anything left is definition or other. An AI Mode question takes the first class it matches, in this order: what percentage|how much|how many|what is a good cost|match rate|worth the price asks for a number; a string starting which or what is the best, or containing \bbest\b, asks which vendor; a string starting how asks how to; a string starting what is, what are, is , do , does or can asks what something is; anything left is other. Matching is on the lowercased string.

Partner suppliers are absent by construction, which is a sampling decision and not a measurement. Cleanlist buys data from Wiza, Hunter, Prospeo, Findymail, AnyMailFinder, Datagma, Icypeas, LeadMagic, Crustdata, Lusha, ZeroBounce and Emailable. We do not publish a shortlist that ranks a company we buy data from, so none of them is in the brand vocabulary and none can score. Their names do appear in the answers Google wrote: Lusha is named in the AI Overviews for "lead enrichment tools" and "contact enrichment tools", and zerobounce.net is cited in 16 answers. The naming leaderboard understates the real market by exactly those names, and the domain leaderboard stops at 15th place, the depth at which it stays clear of a supplier.

Cleanlist ran this study, Cleanlist is in it, and Cleanlist loses the metric that matters

We chose the 240 keywords and the 60 questions, ran the collection, wrote the counting code and are publishing the result, so this is a sample of one market rather than a random sample of Google.

The finding that goes against us is the one to carry away. cleanlist.ai is the 7th most-read domain in these 232 AI Overviews, ahead of apollo.io, g2.com, cognism.com and clay.com. The brand is named in 6 of them, 20th on the naming leaderboard. Google reads us seventh and recommends us twentieth. The dataset is published under CC BY 4.0 below, so every count can be recomputed without us, and the first pass of our own ranking table did not survive that recomputation.

How often does Google write an answer above the results on B2B data buying queries?

On 232 of the 240 keywords, which is 96.7%. All 240 pulls resolved with no errors, so the attempt base and the observation base are the same 240 and there is no gap to reconcile. The 8 keywords with no AI Overview on September 7, 2026 were "hubspot data enrichment", "instantly pricing", "leadiq pricing", "smartlead pricing", "people data labs pricing", "lemlist pricing", "uplead pricing" and "b2b data marketplace". Six of those eight are vendor pricing queries, which is the one corner of this category where Google still hands the buyer ten blue links reasonably often. Everywhere else the answer is written before your page gets a vote, so the shortlist above the results is the category's front door and the names inside it are the shelf.

Why did all 232 answers come back readable when our September 1 study could read only 73 of 196?

Because the collector set the API's asynchronous AI Overview flag this time and did not last time. Google renders the answer body after the initial SERP response, so a collector reading only the first response detects the element and gets no text. Our September 1 study hit exactly that wall: 196 AI Overviews, 73 readable, 37.2%. Almost every published AI Overview study in this category has the same hole, because the flag is off by default.

With the flag set, 162 of the 232 answers arrived on the first response and 70 required the second fetch, so 30.2% of this corpus is text the September 1 method would have discarded as a stub. The shortest body of all 232 is 1,225 characters and there are zero empty bodies. That is the reason to trust these rates more than the last set: the denominator is the whole population of answers rather than the third of it that happened to load in time.

How many vendors does Google put on the shortlist?

A median of 3 in an AI Overview and a median of 4 in AI Mode. Across all 232 AI Overviews the mean is 2.73, the maximum is 8, and 29 answers (12.5%) name nobody. Restricted to the 203 that name at least one vendor, the median is still 3, the mean rises to 3.12 and the minimum is 1. AI Mode writes a longer list: 56 of 60 answers name somebody, the mean is 4.03, and "what is waterfall enrichment and which tools do it" names 9.

The five AI Overviews naming 8 vendors are the broadest category terms in the set: "b2b data provider", "b2b database providers", "b2b data providers", "b2b data vendors" and "b2b data companies". No answer in this corpus names exactly 7. The four AI Mode answers naming nobody are all process questions, including "what should be in a b2b data provider rfp" and "how do I avoid paying twice for the same enriched contact". Google names vendors when the question asks for vendors and writes a procedure when it does not.

SurfaceAnswersName at least oneName nobodyMedian (all)Mean (all)MaxMedian (naming answers only)
Google AI Overview232203 (87.5%)29 (12.5%)32.7383
Google AI Mode6056 (93.3%)4 (6.7%)44.0394

How long is the shortlist? AI Overviews by number of vendors named, base 232

  • AI Overviews
CategoryAI Overviews
0 vendors29
1 vendor39
2 vendors42
3 vendors44
4 vendors37
5 vendors26
6 vendors10
7 vendors0
8 vendors5
Source: Cleanlist buyer shortlist census, September 7, 2026

Which vendors does Google actually name, and at what rate?

ZoomInfo, in more than half of every AI Overview in the category: 121 of 232 (52.2%) and 45 of 60 AI Mode answers (75.0%). HubSpot takes 75 (32.3%), Cognism 73 (31.5%), Clay 66 (28.4%) and Salesforce 44 (19.0%). Cleanlist is named in 6 (2.6%), 20th of the 28 brands appearing anywhere in these AI Overviews, tied with Coresignal.

Two things in this table are worth pausing on. HubSpot, Salesforce and LinkedIn Sales Navigator are platform incumbents rather than enrichment vendors, and one of those three is named in 90 of 232 AI Overviews (38.8%) and 29 of 60 AI Mode answers (48.3%). On 13 of the 203 AI Overviews that name anybody (6.4%), the CRM you already own is the entire shortlist, including on "lead enrichment", "crm data cleansing" and "crm data quality". Second, the surfaces disagree: 6sense and Bombora are named 21 and 15 times in AI Overviews and zero times in AI Mode, while Kaspr goes the other way, 3 against 10.

VendorAI Overviews naming it (of 232)RateAI Mode answers naming it (of 60)Rate
ZoomInfo12152.2%4575.0%
HubSpot7532.3%2135.0%
Cognism7331.5%2135.0%
Clay6628.4%2541.7%
Salesforce4419.0%1728.3%
Clearbit2812.1%1423.3%
LinkedIn Sales Navigator239.9%1220.0%
UpLead239.9%1423.3%
6sense219.1%00.0%
Apollo198.2%2135.0%
Demandbase177.3%35.0%
Bombora156.5%00.0%
People Data Labs125.2%813.3%
NeverBounce114.7%46.7%
RocketReach114.7%11.7%
Salesloft114.7%35.0%
Smartlead114.7%00.0%
Lead41193.9%46.7%
Clearout83.4%11.7%
Cleanlist62.6%11.7%
Coresignal62.6%11.7%
Amplemarket52.2%11.7%
Bettercontact52.2%23.3%
FullEnrich52.2%35.0%
Kaspr31.3%1016.7%
LeadIQ31.3%11.7%
Seamless.AI10.4%00.0%
Surfe10.4%00.0%
Dropcontact00.0%46.7%
Explorium00.0%11.7%
SalesIntel00.0%35.0%
Skrapp00.0%11.7%

Section 3 of the dataset carries these same per-vendor counts, and all 32 rows were recounted from the per-answer rows in sections 1 and 2. Zero mismatches.

Share of 232 Google AI Overviews naming each vendor

  • Share of 232 AI Overviews
CategoryShare of 232 AI Overviews
ZoomInfo52.2%
HubSpot32.3%
Cognism31.5%
Clay28.4%
Salesforce19%
Clearbit12.1%
LinkedIn Sales Nav9.9%
UpLead9.9%
6sense9.1%
Apollo8.2%
Demandbase7.3%
Bombora6.5%
People Data Labs5.2%
NeverBounce4.7%
RocketReach4.7%
Salesloft4.7%
Smartlead4.7%
Lead4113.9%
Clearout3.4%
Cleanlist2.6%
Source: Cleanlist buyer shortlist census, September 7, 2026. Partner suppliers are excluded by construction.

How much of this survives an identical pull 25 minutes later?

The ranking layer survives almost intact and the bottom of the naming layer does not. We collected the same 240 keywords twice on September 7, 2026, about 25 minutes apart, same location, device, language and endpoint. The second pull exists because the first stored the wrong rank field, and the accident is worth more than the bug cost, because almost no published AI Overview study reports run-to-run variance on the same query set. Section 4 of the dataset carries the whole comparison.

AI Overview presence agreed on 229 of the 240 keywords (98.7%). The rank-1 organic result was identical on 96.3% of keywords and the organic top 3 on 85.8%. Over the 229 keywords carrying an AI Overview in both pulls, the mean Jaccard similarity of the two source lists was 0.852.

Brand naming is the unstable layer. Stability below is the share of answers naming a brand in either pull that named it in both.

BrandPull 1Pull 2Naming stability
ZoomInfo12212188.4%
HubSpot737477.1%
Cognism737394.7%
Clay686688.7%
Salesforce454474.5%
Clearbit232859.4%
Apollo231950.0%
UpLead232384.0%
6sense222195.5%
Demandbase161783.3%
Bombora161593.8%
RocketReach121176.9%
Cleanlist11641.7%
NeverBounce1111100.0%
Lead41110990.0%
Kaspr5360.0%
LeadIQ33100.0%
Seamless.AI11100.0%

One row in that table disagrees with the leaderboard above it. Section 4 records HubSpot's second pull at 74 and sections 2 and 3 count 75. The volatility script and the census analyser were written separately and we are publishing the disagreement rather than reconciling it quietly; every other brand agrees across the two files.

Cleanlist is the least stable brand in the table at 41.7%, the lowest of the 18. Our own naming count moved from 11 to 6 across two pulls 25 minutes apart, so the honest report of this baseline is a range, 6 to 11 of roughly 230 answers, and every Cleanlist naming rate on this page is the bottom of it. Of the 12 answers that named us in either pull, 5 named us in both, which is closer to a coin flip than to a measurement.

The column does not simply punish small brands. NeverBounce repeats 11 namings exactly and LeadIQ repeats 3, while Clearbit gains 5 between pulls and Apollo loses 4, and the mean stability of the nine largest brands here (79.1%) is slightly lower than that of the nine smallest (82.9%). What the table does show is that the five largest brands all land between 74.5% and 94.7%, that nothing in it is deterministic, and that the four rows under 65% are Cleanlist at 41.7%, Apollo at 50.0%, Clearbit at 59.4% and Kaspr at 60.0%. The target for a challenger is therefore not one naming but the level of corroboration that stops a name flickering, and any re-measure of this corpus has to run at least twice to see whether it moved.

Which domains does Google actually read?

YouTube, on 107 of the 232 AI Overviews (46.1%), 35 answers clear of second place. Then pipeline.zoominfo.com 72 (31.0%), reddit.com 55 (23.7%), autobound.ai 42 (18.1%), syncgtm.com 38 (16.4%), demandbase.com 34 (14.7%) and cleanlist.ai 32 (13.8%).

Citation here is broad and shallow. There are 1,626 domain-and-answer citation pairs across 408 distinct hostnames, and 204 of those hostnames (50.0%) are cited exactly once in the whole corpus. The top 10 hold 464 of the 1,626 pairs, which is 28.5%, so there is a modest head and a very long tail. Every one of the 232 answers cites something: the per-answer count has a median of 7, a mean of 7.01 and a maximum of 16. AI Mode is stingier, with 183 pairs across 80 hostnames, a median of 3 per answer, and 6 of the 60 answers citing nothing. A domain can enter this field on one good page, and almost none ever move from the tail to the head.

RankHostnameAI Overviews citing itRate of 232
1youtube.com10746.1%
2pipeline.zoominfo.com7231.0%
3reddit.com5523.7%
4autobound.ai4218.1%
5syncgtm.com3816.4%
6demandbase.com3414.7%
7cleanlist.ai3213.8%
8apollo.io3113.4%
8g2.com3113.4%
10cognism.com229.5%
10derrick-app.com229.5%
10sparkle.io229.5%
13clay.com219.1%
13salesmotion.io219.1%
15enginy.ai177.3%

Ranks are competition ranks and ties are broken alphabetically. Two further hostnames, salesforce.com and salesgenie.com, also sit on 17 citations and share 15th place.

Does being cited get you named?

No. Across the 232 AI Overviews there are 633 naming events, and in 458 of them (72.4%) the named vendor's own domain is cited nowhere in that answer. Only 175 namings (27.6%) sit in an answer that cites the named vendor. Clearbit is the clearest case: named 28 times in a corpus where clearbit.com never appears as a source at all.

The source-to-mention ratio is a vendor's namings divided by the number of AI Overviews citing its domain under the roll-up rule. Above 1.0, Google names you more often than it reads you, so other people's pages are carrying your name. Below 1.0, Google reads your pages and recommends somebody else. Cleanlist scores 0.19, which is 6 namings against 32 citations. HubSpot scores 4.41, a 23.2x gap, though HubSpot is a platform incumbent and not a like-for-like comparison. Of the 22 vendors clearing 5 citations, Cleanlist ranks 5th from the bottom. Six vendors are read and never named, and two of those six are detector artifacts described below.

VendorDomainAI Overviews citing itAI Overviews naming itRatioNamed with own domain citedNamed withoutCited without being named
Salesloftsalesloft.com11111.001100
Bomborabombora.com2157.501141
RocketReachrocketreach.co2115.50290
HubSpothubspot.com17754.4114613
Smartleadsmartlead.ai3113.67380
Cognismcognism.com22733.3296413
Clayclay.com22663.0016506
NeverBounceneverbounce.com4112.75470
People Data Labspeopledatalabs.com5122.40570
6sense6sense.com9212.336153
Salesforcesalesforce.com19442.3283611
LinkedIn Sales Navigatorlinkedin.com11232.092219
Lead411lead411.com591.80540
UpLeaduplead.com13231.7732010
Bettercontactbettercontact.rocks351.67320
ZoomInfozoominfo.com741211.64487326
Coresignalcoresignal.com661.00511
Amplemarketamplemarket.com650.83412
Clearoutclearout.io1180.73714
FullEnrichfullenrich.com950.56415
Apolloapollo.io37190.5181129
Demandbasedemandbase.com34170.5010724
Seamless.AIseamless.ai210.50101
Kasprkaspr.io830.38127
Cleanlistcleanlist.ai3260.195127
Snov.iosnov.io1600.000016
Instantlyinstantly.ai1300.000013
Skrappskrapp.io900.00009
SalesIntelsalesintel.io700.00007
Persana AIpersana.ai400.00004
Exploriumexplorium.ai100.00001
Clearbitclearbit.com028no ratio0280
LeadIQleadiq.com03no ratio030
Surfesurfe.com01no ratio010

Ratios built on fewer than 5 citations are unstable. Salesloft's 11.00 rests on one citation and Bombora's 7.50 on two, so read the top of that table as directional. Only 22 vendors clear 5 citations.

Who carries a vendor's name into the answer?

For most vendors, somebody else does. The share of a vendor's namings sitting in an answer that cites its own domain runs from 0.0% (Clearbit, named 28 times with no clearbit.com citation anywhere) to 87.5% (Clearout, 7 of its 8). Cleanlist ties Coresignal for second at 83.3%, 5 of 6, and for us it is a weakness: our name reaches the shortlist almost only when Google is already reading one of our pages, and there are 200 answers where it is not. ZoomInfo's 39.7% is the interesting middle, with 73 of its 121 namings arriving in answers built out of other people's writing, while Clay at 24.2% and Cognism at 12.3% are carried by third parties.

VendorAI Overviews naming itOf those, citing its own domainSelf-sourced share
Clearout8787.5%
Cleanlist6583.3%
Coresignal6583.3%
Amplemarket5480.0%
FullEnrich5480.0%
Bettercontact5360.0%
Demandbase171058.8%
Lead4119555.6%
Apollo19842.1%
People Data Labs12541.7%
ZoomInfo1214839.7%
NeverBounce11436.4%
6sense21628.6%
Smartlead11327.3%
Clay661624.2%
HubSpot751418.7%
Salesforce44818.2%
RocketReach11218.2%
UpLead23313.0%
Cognism73912.3%
Salesloft1119.1%
LinkedIn Sales Navigator2328.7%
Bombora1516.7%
Clearbit2800.0%

That table covers the 24 vendors named 5 or more times.

Does ranking in the organic top ten get you named?

No, and in this corpus it does nothing at all. Cleanlist ranks organic top 10 on 16 of the 240 keywords, all 16 of them carried a readable AI Overview, and Google named us in zero of them. On the other 216 readable answers we are named 6 times (2.8%), so the top-10 group performs worse than the rest of the set, not better. Citation is the condition that predicts naming: on the 32 answers citing cleanlist.ai we are named 5 times (15.6%), and on the 200 that do not, once (0.50%), a factor of 31. That single exception is a false positive described below.

The top-10 set and the named set do not intersect at all. Only 2 of our 6 namings carry a top-20 rank, at positions 15 and 20, and the median organic position across the 5 named keywords that carry a rank at all is 23. Our best organic position in the whole set is 2, on "rocketreach alternative", and that answer does not name us. The same is true of "data enrichment tools" at 5 and "apollo competitors" at 10. Page one and the sentence above page one are close to unrelated.

ConditionKeywordsReadable AI OverviewsNamedRate
Cleanlist ranks organic top 34400.0%
Cleanlist ranks organic top 10161600.0%
Cleanlist ranks organic top 20454524.4%
Cleanlist ranks anywhere in the top 50918955.6%
Cleanlist absent from the top 5014914310.7%
cleanlist.ai cited in the answern/a32515.6%
cleanlist.ai not cited in the answern/a20010.50%

"Absent from the top 50" means absent from the first 50 organic results at the depth we pulled, not absent from Google. All positions are rank_group.

Which of the 6 answers naming Cleanlist actually recommend it?

Five of them. We hand-classified all 6 answer bodies: 5 are genuine shortlist placements and 1 is a false positive where the detector matched the ordinary phrase "clean list" in a sentence about email hygiene. Hand-audited, our figure is 5 of 232 (2.2%) against the 2.6% the automated rule produces.

Both numbers belong here and they are not interchangeable. 2.6% comes from the same automated rule that produced every competitor number on this page, so it is the one to compare against ZoomInfo's 52.2% or Clay's 28.4%. 2.2% is stricter than anything anyone else here has been held to. The audit cannot be recomputed from the CSV, which carries no answer text, so the sentences are printed for a reader to judge.

KeywordOrganic rankCites cleanlist.aiOther vendors namedClassification
data enrichment platform15yesZoomInfo, Clay, Cognismshortlist placement, name inside a link label
clay alternative20yesZoomInfo, Clayshortlist placement
clay alternatives23yesZoomInfo, Clayshortlist placement
lead enrichment tools24yesApollo, ZoomInfo, Clay, Cognism, FullEnrichshortlist placement
clay competitors30yesZoomInfo, Clayshortlist placement
bulk email verificationabsentnononefalse positive on the phrase "clean list"

The five placements read like this. On "clay alternative" and "clay alternatives", which returned the same body:

Cleanlist: Best for pre-built waterfall data enrichment without paying separate extra data credits.

On "lead enrichment tools":

Cleanlist: Best for raw data accuracy, using a waterfall approach that yields high verified email and phone match rates starting at $79/month.

On "clay competitors":

Direct Clay-Style Replacements: SyncGTM and Cleanlist offer multi-provider waterfall enrichment and AI agents with lower-cost pricing plans and independent column refreshing to save credits.

On "data enrichment platform", where the brand name reaches the reader only as the label of a link back to our own page, which is the anchor-text problem described below:

Cleanlist : Runs bulk list cleaning and waterfall enrichment across multiple providers starting at $79 per month.

And the false positive is an ordinary English sentence about a workflow step:

Export: You download the clean list to protect your sender score and stop emails from bouncing.

What kind of question gets a vendor named?

Questions that ask which vendor to buy get the longest shortlist, and questions about price get almost none. In AI Mode, the 36 questions asking which vendor or what is best average 4.56 names and all 36 name somebody. In AI Overviews, the 17 pricing keywords average 1.00 names, the lowest of any class, against 3.79 for "best X". The engine writes a shopping list when asked to shop and a definition when asked what something costs.

One observation here is the most useful thing on the page, and it rests on a single row. The only AI Mode answer of 60 that names Cleanlist is "what percentage of b2b contact data goes stale each year", a question asking for a number. Of the 5 AI Mode questions asking for a figure, 1 names us; of the other 55, none does. Fisher exact on that 2x2 gives a two-sided p of 0.083, so it is not significant and should be read as a hypothesis the corpus is consistent with rather than a law.

It is worth being precise about what that naming is. Google wrote:

You can review further insights on database erosion via ZoomInfo's B2B Data Decay Guide or Cleanlist's Data Decay Statistics.

That is a credit line rather than a shortlist placement, and Cleanlist has zero shortlist placements across 60 AI Mode answers. cleanlist.ai is cited in 5 of the 60 AI Mode answers. One of the five names us. Of the other four, three recommend other vendors and the fourth names no vendor at all. The cleanest case is "best way to enrich hubspot contacts automatically", where our page supplies the opening sentence and HubSpot gets the recommendation:

Automating data enrichment for HubSpot contacts keeps your CRM clean, shortens form fields for better conversion, and gives your sales team instant context (like company size, tech stack, and direct dials).

That sentence is Google's paraphrase of the page it read, not a Cleanlist capability claim, and Cleanlist does not sell technographic data.

Stated as carefully as the evidence allows: publishing a checkable number gets your page read on the questions that ask for numbers, and being read is a weak condition for being named. Whether the number converts a citation into a naming is a claim n=1 cannot support.

Question typeQuestionsNaming at least oneMedian vendorsMean vendorsMaxNaming Cleanlist
asks which vendor / what is best363644.5680
asks how to do something11822.5580
asks for a number5543.2051
asks what something is4455.7590
other433.52.7540
Keyword typeKeywordsNaming at least oneShare naming somebodyMedian vendorsMean vendorsMax
category / tools816681.5%33.158
alternatives / competitors595796.6%32.955
definition or other484083.3%22.256
pricing171482.4%11.002
best X1414100.0%43.796
comparison (vs)10990.0%22.004
how to33100.0%22.002

Both class tables use the ordered rules printed in the method box, so a reader can rebuild them from column one of the dataset.

What does the shortlist sentence look like?

Four parts in a fixed order: the brand as the grammatical subject, "Best for" plus a narrow segment, a mechanism containing a number, and an exact price with a period. The literal string "Best for" appears in 92 of the 232 AI Overviews (39.7%), a dollar figure in 93 (40.1%), and both in 42 (18.1%). The full construction of a bolded brand name immediately followed by "Best for", which is the regex \*\*[^*\n]{1,60}\*\*\s*:?\s*Best for over the answer markdown, appears in 66 (28.4%). In AI Mode, 15 of 60 answers (25.0%) quote prices, 109 dollar figures between them. Those prevalence figures come from a regex pass over the collected answer bodies rather than from the CSV, which carries no answer text.

Here is the shape, verbatim, from three different keywords:

Cleanlist: Best for raw data accuracy, using a waterfall approach that yields high verified email and phone match rates starting at $79/month.

Clay: Best for custom enrichment pipelines. It combines 75+ data sources using waterfall logic (checking multiple providers until it finds a verified match) from around $167/month.

Apollo.io: Best for all-in-one prospecting and multichannel sequences, starting with a limited free tier and paid plans from $49/month.

Google quoted Clay at $167 a month, Clay's real published price. Price is the field the engines quote and the field vendors most often get wrong on each other's pages. In AI Mode the same content arrives as a table with a "Best For" column and a "Pricing Style" column. A category page carrying no compact sentence in this shape under each vendor heading gives the engine nothing to lift.

How concentrated is the naming?

The top three names take 269 of the 633 AI Overview naming events, which is 42.5%. Add the fourth and ZoomInfo 121, HubSpot 75, Cognism 73 and Clay 66 come to 335 of 633, or 52.9%. Four brands take half of every vendor name Google writes in this category. The top five hold 59.9% and the top ten hold 77.9%, which leaves 18 named brands sharing the remaining 22.1%.

AI Mode is flatter: ZoomInfo 45, Clay 25 and Cognism 21 give 91 of 242 naming events (37.6%), with a three-way tie at 21 between Cognism, Apollo and HubSpot. Across both surfaces there are 875 naming events over 292 answers, and ZoomInfo, HubSpot and Cognism hold 40.7%. Cleanlist holds 7 (0.8%), in a three-way tie for 22nd of 32 with Bettercontact and Coresignal.

For a challenger that means a three-to-four name list where four brands are effectively spoken for, and the contest is over the remaining slot or two on narrower keywords rather than on "b2b data providers".

What did this study get wrong?

Six things, four of them corrections we made during analysis and two of them limits of the detector that we are publishing rather than hiding.

1. Our August detector counted URLs as namings, and we published 17.3% when the truth was 9.1%. A brand regex ran over answer text that still held the inline source URLs, so cleanlist matched inside https://www.cleanlist.ai/blog/... and a citation scored as a naming. That one bug roughly doubled our own visibility number. This collection deletes every URL, link target and bare domain first, which is why the 2.6% here is smaller and correct.

2. Anchor text survives that strip, which is the next layer of the same bug. A brand name inside a markdown link label reads as prose to the detector, and one of Cleanlist's 6 namings, the "data enrichment platform" row printed above, is exactly that. Across the corpus the anchor-only share runs from 0% for Clearbit and 13.2% for ZoomInfo, through 14.3% for Cleanlist, up to 62.5% for Clearout, 80% for FullEnrich and 100% for Bettercontact. That pass omits the published homonym guard, so read it as a shape rather than exact values. It inflates some vendors more than others and is the next thing to fix.

3. Our first collector stored the wrong rank field, and every organic position in our internal notes was understated. DataForSEO numbers rank_absolute across every block on the page, so an AI Overview plus a People Also Ask above the fold pushes the same result from organic position 5 to absolute position 7, and 232 of these 240 SERPs carry an AI Overview. The internal ranking table read "top 3: 1 keyword", "top 10: 14" and "top 20: 34" from rank_absolute. Recomputed from rank_group in the published CSV: 4, 16 and 45. Every position on this page is rank_group. The re-collection that produced the corrected field is also the second pull in the volatility section, which is the only reason we can report run-to-run variance at all.

4. Some vocabulary entries cannot be scored at all, and their 0.00 ratios are artifacts. The URL strip that fixes bug 1 also destroys brand names written in domain form. Instantly's only pattern is instantly.ai, and Adapt.io's and Ocean.io's are domain-form too, so a bare .ai or .io domain is deleted before the match runs and the brand can never score. Snov.io carries a prose pattern as well and still scores zero, because every appearance of it in this corpus is the domain. In the published CSV, snov.io is cited in 16 answers and instantly.ai in 13, and both are named in none. Those two 0.00 rows are measurement failures. Skrapp's 9 citations and SalesIntel's 7 with no naming are real.

5. Only Cleanlist's rows were hand-audited, which biases our audited number down against every competitor here. Compare 2.6% with 2.6%, not 2.2% with 2.6%.

6. The ratios here are not comparable to our September 1 study. Cleanlist scored 0.14 there and 0.19 here, on a different query set, vocabulary and answer base. That difference is a method change and nobody should present it as improvement.

What are the limitations of this study?

  1. We are a vendor in our own study, and we chose the queries, ran the collection, wrote the code and published the result.
  2. Partner suppliers are absent by construction. Twelve companies Cleanlist buys data from are excluded from the vocabulary, so the naming leaderboard understates the real market by exactly those names. That is a sampling decision, not a measurement.
  3. The detector uses a fixed 42-name vocabulary and only 32 appear here, so any vendor outside it scores zero by construction. Vendors Google names in these answers that it cannot see include SyncGTM, DemandTools by Validity, OpenRefine and Saleshandy.
  4. Six ambiguous names score only inside a 440-character product-context window, and that filter can undercount. Apollo, with 19 namings against 37 citations of apollo.io, is the row most likely to be depressed by it.
  5. Several vocabulary entries are unmeasurable because of the URL strip, and anchor text inflates namings unevenly by vendor.
  6. Only our own rows were hand-audited.
  7. The "numbers get you named" observation is n=1 with a Fisher exact two-sided p of 0.083, and the naming it produced is a credit line rather than a shortlist placement.
  8. HubSpot, Salesforce and LinkedIn Sales Navigator are platform incumbents. They carry 38.8% of the AI Overview shortlists, and HubSpot's 4.41 is the best stable ratio here without being a like-for-like comparison.
  9. Ratios built on fewer than 5 citations are unstable, and only 22 vendors clear that bar. Small classes carry wide error bars too: the keyword-type table rests on as few as 3 answers.
  10. One collection day, one country, one device, one engine: September 7, 2026, location_code 2840, English, desktop, google.com, no signed-in profile, depth 50. These answers are personalised and non-deterministic. Two identical pulls 25 minutes apart agreed on AI Overview presence 98.7% of the time and on the rank-1 organic result 96.3% of the time, and disagreed heavily on brand naming at the bottom of the leaderboard, where Cleanlist's own stability was 41.7%. Every naming count here is one draw from a distribution, and single-brand counts under about 20 should be read as ranges.
  11. "Absent" means absent from the top 50, not absent from Google.
  12. Three passages here are not recomputable from the published CSV, which carries no answer text: the anchor-text shares, the sentence-shape prevalence figures and the hand audit of Cleanlist's 6 rows. Each is labelled where it appears. Everything else divides columns in the published file.

Where can I download the dataset, and how would somebody reproduce it?

The full open dataset is at /data/buyer-shortlist-census-2026-09.csv under CC BY 4.0. Attribute it to Cleanlist and link back to this page. It is one file with four sections. Sections 1, 2 and 3 each open with a marker row whose first cell begins ## SECTION, then a header row. Section 4 opens the same way but holds two blocks separated by a blank line, and the second of those two blocks starts straight at its header row with no marker. So splitting the file on blank lines gives five blocks, not four, and the fifth block's first row is a header rather than data.

Section 1, 60 rows, one per AI Mode question. query; answer_chars; vendors_named, pipe-separated; vendor_count; names_cleanlist; cited_domains, pipe-separated, lowercased, www. stripped; cited_domain_count; cites_cleanlist; prices_quoted_count; prices_quoted.

Section 2, 240 rows, one per buyer keyword. keyword; has_ai_overview; ai_overview_body_readable; vendors_named; names_cleanlist; cited_domains; cites_cleanlist; cleanlist_organic_rank, the rank_group position, blank when absent from the top 50; cleanlist_url; rank1_domain; rank1_format; top3.

Section 3, 32 rows, one per brand. brand; ai_mode_answers_naming; ai_mode_naming_rate_pct_of_60; ai_overview_answers_naming; ai_overview_naming_rate_pct_of_232.

Section 4, run-to-run volatility, two blocks. First an eight-row metric,value block of agreement statistics across the two pulls. Then, after a blank line and with no marker row, an 18-row brand block: brand; pull1_answers_naming; pull2_answers_naming; named_in_both; named_in_either; naming_stability_pct. Sections 1 to 3 are pull 2.

Recompute the headlines. Coverage is section 2 rows with has_ai_overview equal to 1 (232) over all 240 rows, and the readable base is rows with ai_overview_body_readable equal to 1, also 232. Shortlist length is the length of the pipe-split vendors_named field over the 232 readable rows, then over the 203 naming somebody. The naming leaderboard counts rows whose split vendors_named contains each brand and must reproduce section 3 exactly; it did on all 32. The domain leaderboard splits cited_domains and counts each hostname once per row. The source-to-mention ratio divides namings by the rows whose cited_domains holds the vendor's registrable domain or a subdomain. For the ranking table, cleanlist_organic_rank non-blank and at most 10 gives 16 keywords and names_cleanlist equal to 1 gives 6, and those two sets do not intersect.

Re-collect it without us. Take column one of section 2 as the keyword list and column one of section 1 as the question list. Call any live Google SERP API with location_code 2840, English, desktop, depth 50, and set the asynchronous AI Overview flag, the one thing that makes this corpus different from most published AI Overview studies. On September 7, 2026, 70 of our 232 answers needed that second fetch and all 232 came back with a body of at least 1,225 characters. Read organic positions from the API's rank_group field. For AI Mode, use a Google AI Mode endpoint with the same location and language. For each answer, delete every URL, link target and bare domain first, then match brand names on word boundaries against the lowercased remainder, with a 440-character product-context window for the six ambiguous names. Count a citation once per hostname per answer. Run the whole query set at least twice in the same session, because section 4 shows that one pull is not enough to pin a brand-naming count. Your absolute counts will differ, because these answers are regenerated per request. What should survive is the shape: near-total coverage on buyer queries, a shortlist of three to four names, a top-heavy distribution where four brands take half of everything, a source list that disagrees with the brand list, and a stable ranking layer sitting under an unstable naming layer.

What should a GTM team do with this?

Four things this corpus supports, in the order we are doing them.

Measure naming separately from citation, and strip URLs before you count. Most AI visibility dashboards match your brand against answer text that still holds the source links, which scores every citation as a mention. That bug roughly doubled our own published number in August. Run both metrics, divide by the answers that returned a body, and publish the ratio. Under 1.0 means Google reads you and recommends somebody else.

Stop treating a top-10 ranking as the goal on category keywords. Here it did nothing measurable: we rank organic top 10 on 16 of these keywords and Google named us in none of the 16, against 2.8% on the 216 answers where we do not rank top 10. Being cited is the condition that moves, from 0.50% to 15.6%. We hold position 2 on "rocketreach alternative", our best rank in the whole set, and are not named in that answer. Rank to reach the source pool, then work on what gets you into the sentence.

Write the sentence Google emits. Brand as the subject, "Best for" plus a narrow segment, a mechanism containing a number, then an exact price. Some form of "Best for" appears in 39.7% of these answers and a dollar figure in 40.1%. A comparison page with the price missing or wrong hands the slot to whoever got it right.

Work on the self-sourced share, not just the citation count. 72.4% of all namings here arrive in answers that do not cite the named vendor. Ours run the other way: 5 of our 6 namings need one of our own pages in the source list first, one of the most fragile positions on the leaderboard, because our name then reaches a buyer only on answers Google built out of our writing. It is also why our naming count halved between two pulls 25 minutes apart. The fix is other people's pages, and it is slower than publishing another one of ours.

We will re-run this exact corpus, twice in the same session for the reason section 4 gives, and publish whether any of it moved, including if it did not.

References & Sources

  1. [1]
  2. [2]
  3. [3]
  4. [4]

Enrich your own contacts free for 14 days

Get verified emails and direct dials back on your own contacts and export them to your CRM or sending tool. 1 credit per verified email, 10 per direct dial. 250 credits, 3 seats, 14 days. No card required.

Start free trial

250 credits, 3 seats, 14 days. No card required.

Run it on the list you already have.

Cleanlist puts one lookup through 25+ providers and stops at the first source that returns. Search is free and unlimited on every plan. A verified work email is 1 credit, a direct dial is 10, both together are 11, and a miss costs nothing at all.

14-day Scale trial: 250 credits, 3 seats, no card, every feature except the public API and MCP. The Free plan stays at 30 credits a month after that.