Who Owns Page One in B2B Data: 240 Buyer SERPs, 2,400 Organic Slots

240 buyer keywords, Google US desktop, September 7 2026. No page format holds more than 54 of 240 rank-1 slots, and reddit.com holds 174 of the 2,400 organic slots.

Victor Paraschiv

Victor Paraschiv

Co-Founder & CMO

27 min read

Cleanlist censused 240 B2B data buying keywords on Google US desktop on September 7, 2026, and classified all 2,400 organic top-10 slots. No page format owns this category: the largest rank-1 format holds 54 of 240 SERPs, which is 22.5%. One domain does. reddit.com holds 174 of the 2,400 slots, more than the two largest vendor domains combined. Reddit, YouTube and G2 appear on page one of 182 of the 240 SERPs.

Last updated: September 7, 2026. Every SERP in this study was collected on that date.

Method

Collection. DataForSEO SERP API, serp/google/organic/live/advanced, one live pull per keyword. The same 240-keyword set was collected twice on 2026-09-07, once at 08:47 UTC and once at 09:10 UTC. The published dataset and every figure in this post come from the second pull, stamped 2026-09-07 09:10:46 UTC. Search engine google.com, location_code 2840 (United States), language_code en, device desktop, depth 50, asynchronous AI Overview loading enabled. No personalisation and no signed-in profile. 240 keywords attempted, 240 resolved, 0 errors.

Sample frame. 240 buyer-intent keywords in one category: B2B contact data, enrichment, prospecting and list hygiene. They were drawn from a 354-keyword seed universe by ranking on monthly volume multiplied by CPC, with 22 strategic terms forced in. The seed list was filtered so that no company Cleanlist buys data from appears in any keyword string. This is a purposive sample of one market, not a random sample of Google.

Counting rules. A slot is one organic result in positions 1 to 10. Every one of the 240 SERPs returned exactly ten, so the slot denominator is 2,400 with no partial pages. Every slot carries exactly one format label from an eleven-value classifier, published in full below, applied to the URL, the page title and the domain with first match winning. Domain leaderboards are computed on a lowercased domain with a leading www. stripped, and subdomains are kept distinct (pipeline.zoominfo.com and zoominfo.com are separate rows). Section 1's rank1_domain column is the exception: it is printed as the API returned it, so 156 of its 240 values still carry a leading www..

Appearances versus SERPs. A domain can hold two slots on one page. Where this study says appearances it means slots, and where it says SERPs it means distinct keywords. Both are printed for every leaderboard row in this post, because they differ. Section 3 of the published CSV carries the appearance count only.

Volume and price. Google Ads US monthly search volume and top-of-page CPC come from keywords_data/google_ads/search_volume/live on the same location and language. 238 of the 240 keywords carry a volume and 212 carry a CPC. Every volume rate divides by 238 and every CPC rate divides by 212, never by 240.

Fan-out. People Also Ask and related searches were read from the same SERP responses on an 80-keyword subset. 74 returned a People Also Ask block, each with exactly four questions, giving 296 keyword-and-question pairs and 215 unique questions. Fan-out rates divide by 296 or by 74, never by 80.

AI Overviews. An AI Overview was present on 232 of the 240 SERPs (96.7%). The eight without one were hubspot data enrichment, instantly pricing, leadiq pricing, smartlead pricing, people data labs pricing, lemlist pricing, uplead pricing and b2b data marketplace.

What is not in the published CSV. Five kinds of figure in this post cannot be recomputed from the census CSV alone, and each is flagged where it appears: the distinct-SERP count behind any domain's appearance count (Section 3 publishes appearances only), the number of distinct domains outside the published top 60, the wider non-vendor domain classification, the two-pull agreement figures, and the 354-keyword seed total. All five come from the raw SERP responses or the working keyword-volume file. The domain lists and the rebuild recipe for each are printed at the end of this post.

Cleanlist ran this study and Cleanlist loses in it

We chose the keywords, wrote the classifier and are publishing the result, and our own domain is one of the rows. Cleanlist is absent from the top 50 on 149 of the 240 SERPs, holds no number-one position on any of them, and is missing from 66.6% of the monthly search volume this study measured. That table is printed below in full, along with every ranking we do hold. The dataset is published under CC BY 4.0 so the whole thing can be recomputed without us.

What kind of page ranks first on a B2B data buying query?

No single format ranks first often enough to call it the answer. Across 240 buyer SERPs the most common rank-1 format is the vendor category page at 54, which is 22.5%, and the second is the editorial listicle ("14 Best Data Enrichment Tools") at 50, which is 20.8%. Alternatives pages take 37, community threads 30, glossary and definition pages 24, pricing pages 19, blog posts 9, review sites 9, comparison pages 5 and documentation 3. The top four formats together hold 171 of 240, or 71.2%, and every one of the 240 rows carries a label with no blanks. Grouped into three families, a vendor-owned page holds rank 1 on 109 SERPs (45.4%), an editorial ranking page on 92 (38.3%), and a third-party page belonging to nobody selling anything on 39 (16.2%). The practical reading for a content plan is that betting an entire quarter on one page shape loses about four buyer queries in five.

Format holding the number one organic result, 240 B2B data buyer keywords

  • Keywords where it ranks first
CategoryKeywords where it ranks first
Vendor category page54
Listicle50
Alternatives page37
Community thread30
Glossary or definition24
Pricing page19
Blog post9
Review site9
Comparison page5
Docs3
Source: Cleanlist B2B Buyer SERP Format Census, September 7 2026 (n = 240)
Rank-1 formatKeywords where it ranks firstShare of 240FamilyMedian CPC of those keywords
vendor_category_page5422.5%vendor-owned page$41.56 (n=46 priced)
listicle5020.8%editorial ranking page$53.11 (n=43)
alternatives3715.4%editorial ranking page$42.05 (n=37)
community3012.5%third-party non-vendor$47.08 (n=26)
glossary_definition2410.0%vendor-owned page$49.51 (n=22)
pricing197.9%vendor-owned page$14.06 (n=19)
blog93.8%vendor-owned page$20.91 (n=4)
review_site93.8%third-party non-vendor$52.87 (n=9)
comparison52.1%editorial ranking page$40.46 (n=5)
docs31.3%vendor-owned page$36.97 (n=1)
TOTAL240100%vendor-owned 109 / editorial 92 / non-vendor 39$45.80 median across all 212

How does the classifier decide, and where does it get things wrong?

The classifier is a URL-and-title heuristic with first-match-wins ordering, everything lowercased: community if the domain is reddit.com, quora.com or news.ycombinator.com; else video for youtube.com or vimeo.com; else review site for g2.com, capterra.com, getapp.com, trustradius.com, softwareadvice, gartner.com or sourceforge; else docs if the URL matches /docs?/, /api/ or developer.; else glossary if the URL matches /glossary/, /what-is-, /learn/ or /resources/.*/what, or the title starts with "what is"; else comparison on \bvs\b or "versus" in the title or -vs- in the URL; else alternatives on alternatives? in the title or URL; else pricing on \bpricing\b, \bcost\b or \bprice\b in the title; else listicle on a title starting with a digit or containing "best" or "top N"; else blog on /blog/, /post/ or /article in the URL; else vendor_category_page.

Two things follow from those rules and both cost accuracy. alternatives is a page shape rather than a publisher type, so it fires on a vendor's own /alternatives/ directory as readily as on a third-party roundup. And vendor_category_page is the unmatched fall-through bucket, so it absorbs everything the rules miss and is the least trustworthy label in the file. A minority of the 2,400 rows carry the wrong label. The rules are printed here so that a reader can re-run them and disagree with specific rows.

How much of page one is not a vendor at all?

Between one slot in seven and one in six, depending on where you draw the line, and the line is a judgement call rather than a fact. Under this study's own classifier, community, review-site and video pages hold 331 of the 2,400 organic top-10 slots, which is 13.79%, and at least one of them appears on 188 of the 240 SERPs (78.3%), with a mean of 1.38 such slots per SERP, a median of 1 and a maximum of 6. Under a wider domain list that adds forums, social networks, directories and aggregators, printed in full at the end of this post, the non-vendor total rises to 430 slots (17.92%) present on 219 of 240 SERPs (91.3%), split as community 211, review sites 105, directories and aggregators 61, and video 53. Everything else, 1,970 slots or 82.08%, is a vendor or publisher site. That wider split is computed from the raw SERP responses rather than from the published CSV, which carries format labels rather than a domain on every slot.

reddit.com alone accounts for 174 slots across 159 distinct SERPs, which is 66.2% of the corpus, and it holds an organic top-3 position on 85 SERPs (35.4%, or 11.81% of all 720 top-3 slots). The 159 is the one number in that sentence that needs the raw responses rather than the CSV. A non-vendor publisher holds the number one result outright on 39 of 240 SERPs (16.2%) under this study's own classifier: reddit.com 30, gartner.com 7, g2.com 2. The classifier's review-site list omits trustpilot.com, which holds rank 1 on five more and is labelled vendor_category_page instead, so on a domain judgement rather than on the published label the count is 44 (18.3%). Both are recomputable from Section 1: the 39 from rank1_format, the 44 from rank1_domain.

FormatSlotsShare of 2,400SERPs where it holds rank 1
listicle54222.58%50
vendor_category_page48320.12%54
alternatives42417.67%37
community1857.71%30
blog1807.50%9
pricing1586.58%19
glossary_definition1335.54%24
comparison1164.83%5
review_site933.88%9
video532.21%0
docs331.38%3
Non-vendor subtotal33113.79%39

YouTube is the odd row. It holds 53 slots on 48 SERPs, and it holds an organic top-3 position on zero of the 240. It is on page one often and never at the top of it.

Which domains own this category?

497 distinct domains hold at least one of the 2,400 slots, and 233 of them (46.9%) hold exactly one. Those two counts need the raw responses; everything else in this paragraph comes from the published leaderboard. The head is fat and the tail is very long: the top 5 domains hold 442 slots (18.4%), the top 10 hold 647 (27.0%), and the top 20 hold 926 (38.6%). The published leaderboard is truncated to 60 domains, which cover 1,456 slots or 60.7% of the corpus, so anyone computing concentration from the published file alone will understate the tail.

Rank 1 is even less concentrated than page one. 90 distinct domains hold the number one result across the 240 SERPs and 56 of them (62.2%) hold it exactly once. reddit.com leads with 30 rank-1 slots (12.5% of 240), then zapier.com 18, demandbase.com 15, saleshandy.com 14 and rb2b.com 9.

Top 15 domains by organic top-10 appearances across 240 buyer SERPs

  • Top-10 slot appearances
CategoryTop-10 slot appearances
reddit.com174
pipeline.zoominfo.com104
cognism.com60
youtube.com53
snov.io51
demandbase.com49
g2.com45
sparkle.io39
saleshandy.com36
salesforge.ai36
kaspr.io35
zapier.com34
salesgenie.com33
amplemarket.com32
gartner.com30
Source: Cleanlist B2B Buyer SERP Format Census, September 7 2026 (2,400 slots)
#DomainTop-10 appearancesShare of 2,400 slotsDistinct SERPs (of 240)What it is
1reddit.com1747.25%159 (66.2%)forum
2pipeline.zoominfo.com1044.33%102 (42.5%)vendor content hub
3cognism.com602.50%58 (24.2%)B2B data vendor
4youtube.com532.21%48 (20.0%)video platform
5snov.io512.13%37 (15.4%)B2B data vendor
6demandbase.com492.04%46 (19.2%)B2B data vendor
7g2.com451.88%45 (18.8%)review marketplace
8sparkle.io391.63%37 (15.4%)GTM software vendor
9saleshandy.com361.50%36 (15.0%)cold-email vendor
10salesforge.ai361.50%36 (15.0%)cold-email vendor
11kaspr.io351.46%32 (13.3%)B2B data vendor
12zapier.com341.42%33 (13.8%)automation platform ranking as a publisher
13salesgenie.com331.38%33 (13.8%)B2B data vendor
14amplemarket.com321.33%28 (11.7%)GTM software vendor
15gartner.com301.25%30 (12.5%)analyst and review site
-cleanlist.ai (for reference)160.67%16 (6.7%)the study's author

Three of the top seven rows are a forum, a video platform and a review marketplace. Together those three hold 272 of the 2,400 slots (11.33%), which the published leaderboard gives directly, and at least one of them is on page one of 182 of the 240 SERPs (75.8%), which needs the raw responses. None of them sells B2B data.

Read this table as a census of which domains appear, ordered by appearance count. It is not a quality ranking and not a competitive ranking. Several domains in the full 60-row published list are companies Cleanlist buys data from, and they are counted by exactly the same rule as everybody else.

Does a high-CPC keyword have a different page one from a low-CPC one?

No. Once one confound is controlled for, the format mix on an expensive keyword is indistinguishable from the format mix on a cheap one. Sorting the 212 priced keywords into equal CPC terciles at boundaries of $31.32 and $57.27, the grouped rank-1 family mix looks like it differs: chi-square 11.71 on 4 degrees of freedom, p = 0.020. That result is driven entirely by a single format. A pricing page holds rank 1 on 15 of the 70 lowest-CPC keywords (21.4%) and on 1 of the 72 highest (1.4%), a two-proportion z of -3.78 and p = 0.0002. The reason is a query-type confound rather than a price effect: 23 of the 212 priced keywords are pricing-intent queries whose own text contains "pricing", "price" or "cost", a pricing page wins rank 1 on 19 of those 23 (82.6%), and their median CPC is $20.75 against $49.09 for everything else. Pricing questions are cheap to buy and are answered by pricing pages.

Remove those 23 keywords and re-run on the remaining 189 in terciles of 63, and the effect disappears: chi-square 6.72 on 4 degrees of freedom, p = 0.151, with no individual format reaching significance. Rank correlations agree. Across all 212 priced keywords, Spearman rho between CPC and the number of non-vendor slots on page one is -0.050, against community slots it is -0.115, and against monthly volume it is -0.138. With pricing keywords removed it is -0.085. The honest finding is a null result with one confound identified. Paying more per click does not buy you a friendlier page one.

CPC bandnCPC rangeMean CPCEditorial @1Vendor-owned @1Non-vendor @1Pricing page @1Mean non-vendor slots
low70$0.89 to $31.24$17.9722 (31.4%)41 (58.6%)7 (10.0%)15 (21.4%)1.54
mid70$31.32 to $56.71$44.3928 (40.0%)26 (37.1%)16 (22.9%)3 (4.3%)1.49
high72$57.27 to $363.60$102.4935 (48.6%)25 (34.7%)12 (16.7%)1 (1.4%)1.39
test, all 212chi-square 11.71, df 4p = 0.020z = -3.78, p = 0.0002Spearman -0.050
control, 189 keywords, pricing intent removed63 / 63 / 6330 / 22 / 3325 / 24 / 208 / 17 / 10n/aSpearman -0.085
test, controlchi-square 6.72, df 4p = 0.151no format significantno format significantno format significantno effect

What does the money look like in this category?

Low volume and very high prices. The 240 censused keywords carry 69,680 US searches a month across the 238 that report a volume, and 215 of those 238 (90.3%) get under 500 searches a month. The median keyword gets 90 searches a month; the mean is 292.8, pulled up by the single largest term, "email verification tools" at 9,900 a month. The wider 354-keyword seed universe carries 70,950 searches a month, a figure that lives in the working keyword-volume file rather than in the published CSV, so the censused set covers 98.2% of the demand this project measured.

Prices go the other way. Across the 212 keywords carrying a CPC, the median top-of-page cost per click is $45.80 and the mean is $55.40, with a 25th percentile of $25.33, a 75th of $70.03 and a 90th of $101.15. 96 of the 212 (45.3%) cost $50 or more per click, 23 (10.8%) cost $100 or more, and 6 (2.8%) cost $200 or more. The most expensive click in the corpus is "best b2b contact database" at $363.60 against 40 searches a month, and reddit.com holds its number one organic result. Buying every click in this corpus at its published CPC would cost about $2.34 million a month, or $28.05 million a year, on 68,770 monthly searches at a volume-weighted mean CPC of $34.00. Treat that as a ceiling and nothing more: it assumes a 100% click share, which nobody gets, on volumes Google models rather than measures.

MeasureValueDenominator
Keywords carrying a monthly volume238of 240
Total US searches a month, censused set69,680238 keywords
Total US searches a month, 354-keyword seed universe70,950283 of 354 with a non-zero volume
Median monthly volume90238
Keywords under 500 searches a month215 (90.3%)238
Largest single keywordemail verification tools, 9,900/mo, CPC $10.08238
Keywords carrying a CPC212of 240
Median top-of-page CPC$45.80212
Mean CPC$55.40212
p25 / p75 / p90 CPC$25.33 / $70.03 / $101.15212
CPC at or above $5096 (45.3%)212
CPC at or above $10023 (10.8%)212
CPC at or above $2006 (2.8%)212
Highest CPC in the corpus$363.60, best b2b contact database, 40 searches a month212
Volume-weighted mean CPC$34.00212 keywords, 68,770 searches
Cost of buying 100% of that click volume$2,337,873 a month, $28,054,477 a yeara ceiling, not a forecast
Google Ads competition ratingLOW 131, MEDIUM 89, HIGH 17237 rated of 240

The twelve most expensive clicks, with the format holding rank 1 on each, taken from the published file rather than from a summary:

KeywordCPC (USD)Monthly US searchesCompetitionFormat at rank 1Domain at rank 1
best b2b contact database$363.6040LOWcommunityreddit.com
top b2b data providers$274.5020LOWlisticledemandbase.com
data enrichment companies$245.66170LOWlisticlezapier.com
demandbase competitors$229.1590MEDIUMreview_sitegartner.com
prospect database$204.00110MEDIUMglossary_definitiondatarade.ai
prospecting database$204.00110MEDIUMvendor_category_pagesalesql.com
phone appending services$188.3370MEDIUMlisticleaccurateappend.com
bombora competitors$156.6220MEDIUMreview_sitegartner.com
crm enrichment$155.16110LOWglossary_definitionclay.com
sales intelligence software$129.19170LOWglossary_definitionibm.com
ai sdr tools$127.87210LOWblogsalesforge.ai
bombora pricing$127.8050MEDIUMvendor_category_pagebombora.com

Two things in that table are worth a marketer's attention. The single most expensive click in the entire category is won organically by a Reddit thread, which means a buyer clicking the $363.60 ad and a buyer clicking the top organic result are reading very different things. And "prospect database" and "prospecting database" both read 110 searches a month at $204.00, which is Google reporting the same bucketed estimate for two variants. Adjacent variants in this file are not independent demand.

What does Google think a buyer wants to know next?

It thinks the buyer wants a definition. Across 296 People Also Ask questions harvested from 74 of these keywords, 177 (59.8%) open with "what is", "what are", "what's" or "what does X mean". How-questions take 42 (14.2%), yes-or-no auxiliaries 41 (13.9%), "who" 14 (4.7%), "which" 12 (4.1%), other "what" forms 7 (2.4%), "why" 2 and "where" 1. Beyond the opening word, 67 of the 296 (22.6%) name a tool, software, platform, provider or vendor, by the regex \b(tools?|software|platforms?|providers?|vendors?)\b applied to the whole question; 53 (17.9%) contain the word "best"; 29 (9.8%) ask about price or cost; and 20 (6.8%) ask about something free. Only 12 of the 296 rows (4.1%) expose the URL Google answered the question from, so this map shows what Google asks and cannot show who it believes.

The repeated questions are the useful part, because a question that recurs across different keywords is a hole in the category's published writing. The most repeated is a price question.

QuestionKeywords it appears underOf 296 pairs
how much does b2b data cost?103.4%
what is the best tool for data enrichment?82.7%
is there a free b2b contact database available?72.4%
what is an example of data enrichment?62.0%
what is the rule of 7 in b2b?51.7%
what are the top 5 data management tools?41.4%
what is the best b2b contact database right now?41.4%
is there anything better than apollo?41.4%
will crm be replaced by ai?31.0%
what is a b2b data provider?31.0%
who are the top b2b data providers?31.0%
what is the most hacked email provider?31.0%
who is apollo's biggest competitor?31.0%
who is zoominfo's biggest competitor?31.0%

Where does Cleanlist rank in its own study?

Badly, and the table below is the least flattering thing on this website. Cleanlist is absent from the top 50 on 149 of the 240 SERPs (62.1%) and appears anywhere in the depth-50 pull on 91 (37.9%). The published cleanlist_rank column is the organic position, so it needs no adjustment: top-3 on 4, top-5 on 8, top-10 on 16, top-20 on 45 and top-30 on 70, and number one on none of the 240. Where it does rank, the median position is 21, the mean is 21.1, the best is 2 and the worst is 46. Weighting by demand instead of by keyword makes it worse: the 148 priced keywords we are absent from carry 46,380 of the 69,680 monthly searches in this corpus, so we are missing from 66.6% of the measured demand. cleanlist.ai holds 16 of the 2,400 slots, which is 0.67%, across 16 distinct SERPs.

CutKeywordsShare of 240
Number one00.0%
Top 341.7%
Top 583.3%
Top 10166.7%
Top 204518.8%
Top 307029.2%
Anywhere in the depth-50 pull9137.9%
Absent from the top 5014962.1%
Median position where present21-
Share of measured monthly volume we are absent from46,380 of 69,68066.6%

Every top-10 ranking we hold in the corpus, in full:

Organic positionKeywordMonthly volumeCPCFormat at rank 1
2rocketreach alternative210$32.39alternatives
3apollo vs zoominfo260$52.41community
3database cleaning services320$35.09community
3rocketreach competitors90$42.05alternatives
4rocketreach alternatives210$32.39alternatives
5clearbit pricing140$31.01vendor_category_page
5data appending320$44.22glossary_definition
5data enrichment tools390$54.93listicle
6crm data enrichment tools20no CPCcommunity
6data enrichment vendors40no CPClisticle
6how to enrich a csv of leadsno volumeno CPCblog
7data enrichment software170$24.24listicle
8ai data enrichment40$27.79glossary_definition
10amplemarket pricing50$9.76pricing
10apollo competitors170$52.48alternatives
10lead list building10no CPCvendor_category_page

Sixteen rankings on 240 buyer queries, none of them first, and the largest keyword in the set gets 390 searches a month. That is what a category SEO programme looks like from the inside when you measure it against the whole buying surface instead of against your own best pages.

What did this study get wrong?

Eight things, all found by recomputing the published file rather than trusting the analysis that produced it. Five are defects in the dataset itself and are stated here because the file is public and somebody will hit them.

DefectWhat the file saysWhat is trueEffect
reddit_in_top10 is computed on the top threesums to 85 of 240Reddit is in the organic top 10 on 159 of 240 (66.2%). 85 is the top-3 countThe column name is wrong. Use the domain leaderboard for a top-10 figure
youtube_in_top10 is computed on the top threesums to 0 of 240YouTube genuinely never holds an organic top-3 slot, and is in the top 10 on 48 of 240The zero is true of the top 3 and misleading as a top-10 flag
Section 3 counts appearances, not SERPsreddit.com 174174 slots across 159 distinct SERPs. Reddit takes two top-10 slots on 15 SERPsAny "appears on N of 240 SERPs" claim built on Section 3 is inflated
Section 3 is truncated to 60 domains60 rows497 domains hold at least one of the 2,400 slots, and the 60 cover 60.7% of themThe long tail is invisible in the published file
Section 2 prints one share rounded up at a half-way pointdocs 1.3%3/240 = 1.25% exactly, which a round-half-to-even convention prints as 1.2%. The other nine shares are correct to one decimal either way, including alternatives at 15.4%Counts are exact; only the printed percentage differs
The related-searches count is double-counted upstreamworking file holds 1,104 rows73 of the 74 keywords return the related-searches block twice. Deduplicated within keyword the total is 564, with 426 unique stringsDo not publish 1,104. This post does not use the figure at all
The fan-out covers 74 keywords, not 8080 attempted6 of the 80 returned no People Also Ask block296 pairs across 74 keywords, exactly 4 per keyword
The format classifier is a heuristicone label per result"clay alternatives" is a genuine homonym SERP about pottery. alternatives fires on a vendor's own /alternatives/ page. vendor_category_page is the fall-through bucketA minority of rows carry the wrong label. The rules are published so they can be re-run

The homonym case is worth showing, because Google's own AI Overview for that query splits the difference and says so. Verbatim, with the answer's bold and code formatting removed: "The phrase clay alternatives can refer either to crafting materials or to data enrichment software, so here are the top options for both categories." A keyword that a category treats as a competitor term is, to Google, half a pottery query.

What are the limitations of this study?

  1. One day, two pulls. One country (location_code 2840), one device (desktop), one search engine, no signed-in profile, one date (2026-09-07). The same 240 keywords were collected twice that night, 23 minutes apart at 08:47 and 09:10 UTC, which measures run-to-run variance directly instead of disclaiming it. The rank-1 organic result was identical on 231 of 240 (96.3%), the whole organic top 3 was identical on 206 of 240 (85.8%), an AI Overview was present on 230 in the first pull and 232 in the second with 229 carrying one in both (98.7%), and the mean Jaccard overlap of AI Overview source lists across those 229 shared answers was 0.852. Every figure in this post is the second pull; the agreement figures come from comparing the two raw response files and are not recomputable from the published CSV. Single positions move on re-collection. The shape does not.
  2. The keyword set is ours. It is a purposive sample of one category, selected by ranking a 354-keyword seed universe on volume multiplied by CPC with 22 strategic terms forced in, and filtered so that no Cleanlist supplier appears in any keyword string. It is deliberately biased toward commercially valuable terms and generalises to no other category without re-collection.
  3. The classifier misclassifies rows. It reads a URL, a title and a domain. It cannot tell a third-party roundup from a vendor's own alternatives directory, and its fall-through bucket is also its second-largest label.
  4. "Non-vendor" is a domain-list decision. 13.79% under the study's own classifier, 17.92% under the wider list. Both lists and both numbers are published so a reader can pick, and neither is a fact about the internet.
  5. The CPC test is cross-sectional. 212 keywords on one day, not an experiment. The grouped effect that looks significant at p = 0.020 collapses to p = 0.151 once 23 pricing-intent keywords are removed. Report it as a null result with a confound, not as an effect.
  6. The $2.34 million figure is a ceiling. It assumes a 100% click share at the published top-of-page CPC on 68,770 monthly searches. Nobody gets a 100% click share, and Google Ads volumes are bucketed and modelled rather than measured.
  7. Adjacent keyword variants are not independent. "prospect database" and "prospecting database" both read 110 a month at $204.00; "sales intelligence platform" and "sales intelligence platforms" both read 1,600 a month at $66.42.
  8. The fan-out is a smaller and shallower corpus. 80 keywords attempted, 74 returning a People Also Ask block, of which 66 are also in the 240-keyword census. Exactly four questions came back per keyword, which is an artifact of the endpoint rather than a fact about how many questions Google holds. Only 12 of the 296 rows expose an answer source, so the map cannot say who Google answers from.
  9. The published top-10 leaderboard is truncated at 60 domains. Concentration computed from it alone understates the tail by 437 domains.
  10. We are a vendor in our own study. Cleanlist chose the keywords, wrote the classifier, ran the collection and is publishing the result, and the rows where we lose are printed above rather than summarised.

Where can I download the dataset, and how would somebody reproduce it?

Two files, both under CC BY 4.0. Attribute them to Cleanlist and link back to this page.

Parsing the census file. It is one CSV with three sections. Split it on a blank line followed by a row beginning ## SECTION; each section carries its own header row. Section 1 is 240 rows, one per keyword. Section 2 is the 10-row rank-1 format distribution. Section 3 is the top 60 domains by top-10 appearance. Sections 2 and 3 are both derivable from Section 1 plus the raw SERPs, and Section 2 reproduces exactly from Section 1's rank1_format column.

Section 1 column dictionary. keyword, the search string. google_ads_monthly_volume_us, blank on 2 rows. cpc_usd, top-of-page cost per click, blank on 28 rows. competition, LOW, MEDIUM or HIGH, blank on 3. ai_overview_present, 1 or 0. rank1_domain, the domain holding the number one organic result. rank1_format, one of the eleven classifier labels. top10_format_mix, pipe-separated format:count pairs summing to 10 on every row. reddit_in_top10 and youtube_in_top10, both of which are top-3 flags despite their names. cleanlist_rank, the organic position (DataForSEO rank_group, never rank_absolute) of the best-ranking cleanlist.ai URL in the depth-50 pull, blank on 149 rows. cleanlist_url, populated on exactly the 91 rows that carry a rank.

Fan-out column dictionary. keyword, question, google_answer_source_url. One section, three columns, 296 rows, four per keyword across 74 keywords, 12 rows carrying a source URL.

Rebuilding every figure. Rank-1 distribution: value counts of rank1_format over all 240 rows. Page-one composition: parse top10_format_mix and sum across all rows, which totals 2,400. Non-vendor share: sum the community, review_site and video counts for 331. Money: sum google_ads_monthly_volume_us over the 238 non-blank rows for 69,680, and take percentiles of cpc_usd over the 212 non-blank rows for a $45.80 median. CPC test: sort the 212 priced rows by CPC, cut into terciles of 70/70/72 at $31.32 and $57.27, cross-tabulate against rank1_format grouped into three families, then repeat after dropping the 23 keywords whose text matches pricing|price|cost and watch the effect disappear. Cleanlist: threshold cleanlist_rank at 3, 5, 10, 20, 30 and 50, and count blanks for the absent figure. Fan-out shapes: first-match-wins regex on the opening word, with the "what is" class defined as ^what\s+(is|are|'s|does .* mean) and the tool-shaped class defined as \b(tools?|software|platforms?|providers?|vendors?)\b anywhere in the question.

Re-collecting from scratch. Take column one of Section 1 as the keyword list. Call any live Google SERP API with search engine google.com, location_code 2840, language_code en, device desktop, depth 50, and asynchronous AI Overview loading enabled. Ours was DataForSEO's serp/google/organic/live/advanced, one task per POST. Volume and CPC come from keywords_data/google_ads/search_volume/live on the same location and language. The fan-out comes from the people_also_ask and related_searches blocks of the same SERP responses, deduplicated within each keyword.

Two corrections to apply if you are checking us. reddit_in_top10 and youtube_in_top10 are computed on the top three, so they sum to 85 and 0 while the organic top-10 counts are 159 and 48. And Section 3 counts appearances rather than distinct SERPs, so reddit.com's 174 slots sit across 159 SERPs. cleanlist_rank needs no correction: it is the organic position (rank_group), not DataForSEO's rank_absolute, which counts every block on the page and therefore reads one or two places lower wherever an AI Overview or a People Also Ask block sits above the first organic result.

The wider non-vendor domain list, in full. The 17.92% figure matches on the lowercased domain with a leading www. stripped. Community: reddit.com, quora.com, news.ycombinator.com, stackoverflow.com, linkedin.com, medium.com, x.com, twitter.com, facebook.com, substack.com, indeed.com, glassdoor.com. Video: youtube.com, vimeo.com. Review: g2.com, capterra.com, getapp.com, trustradius.com, softwareadvice.com, gartner.com, sourceforge.net, trustpilot.com, peerspot.com, saasworthy.com, financesonline.com, selecthub.com, slashdot.org, tekpon.com, softwaresuggest.com, goodfirms.co, clutch.co. Directory and aggregator: crunchbase.com, about.crunchbase.com, datarade.ai, alternativeto.net, saashub.com, producthunt.com, github.com, theresanaiforthat.com, aitools.fyi, zapier.com, wikipedia.org, en.wikipedia.org. Every other domain counts as a vendor or publisher site. The list is a judgement call, and it is printed so a reader can argue with a specific domain rather than with the number.

Expect your absolute positions to differ from ours, because Google moves. What should survive re-collection is the shape: no format above about 23% of rank 1, a forum as the single largest domain, about one page-one slot in seven held by nobody selling anything, and a very long tail of one-hit domains.

What should a GTM team do with this?

Four things this corpus supports, in any category and not only ours.

Census your own page one before you write a content plan. The format distribution here is flat enough that a plan built entirely on listicles, or entirely on category pages, addresses about a fifth of the buying surface. Pull 100 to 250 of your buyer keywords at depth 10, label what holds rank 1, and let the mix decide the ratio of page types you build. It costs a couple of dollars in API calls.

Price your entry against CPC, not against volume. The median keyword here gets 90 searches a month and costs $45.80 a click, and 90.3% of the set is under 500 searches a month. A keyword doing 40 searches a month at $363.60 is worth more than a keyword doing 1,000 at $2. Rank the list on volume multiplied by CPC, and be ready for the answer that the whole category is small and expensive rather than large and cheap.

Treat forums, review sites and video as a distinct surface with its own tactics. Reddit, YouTube and G2 are on page one of 75.8% of these SERPs, and a non-vendor publisher holds the number one result outright on 16.2% of them. No amount of on-site publishing moves those slots. Being present there is a different job with a different owner, and in most B2B categories nobody owns it.

Answer the definitional questions on the page that sells. Google's own fan-out for this category is 59.8% definitional and only 22.6% tool-shaped, and its single most repeated question across different keywords is "how much does b2b data cost?". A category page that answers the definition and the price on the same URL is aimed at what Google is actually asking, and both are cheaper to write than another roundup.

Enrich your own contacts free for 14 days

Get verified emails and direct dials back on your own contacts and export them to your CRM or sending tool. 1 credit per verified email, 10 per direct dial. 250 credits, 3 seats, 14 days. No card required.

Start free trial

250 credits, 3 seats, 14 days. No card required.

Run it on the list you already have.

Cleanlist puts one lookup through 25+ providers and stops at the first source that returns. Search is free and unlimited on every plan. A verified work email is 1 credit, a direct dial is 10, both together are 11, and a miss costs nothing at all.

14-day Scale trial: 250 credits, 3 seats, no card, every feature except the public API and MCP. The Free plan stays at 30 credits a month after that.