33% off, foreverNew customers · code · ends Aug 31
Claim
guidesprimary researchAI searchAEO

AI Visibility Index 2026: Which B2B Data Tools AI Engines Recommend

Cleanlist put 72 B2B buying questions to 5 AI answer engines twice, 4 days apart. 712 answers, 51 tools. 32% of the tools named on Aug 3 were gone by Aug 7. Corrected Aug 7. Open dataset, CC BY 4.0.

Victor Paraschiv

Victor Paraschiv

Co-Founder & COO

August 7, 2026
31 min read
AI Visibility Index 2026: share of voice for B2B data tools across five AI answer engines

Cleanlist ran the AI Visibility Index for B2B data tools: 72 real buying questions, put to five AI answer engines twice, four days apart, producing 712 archived answers that named 51 tools. Apollo is named in 75.6% of answers, ZoomInfo in 57.9%, Clay in 46.5%, Cognism in 38.6%, Lusha in 36.1%, HubSpot in 32.9%, Hunter.io in 29.6%, Salesforce in 23.3%, Clearbit in 20.4% and NeverBounce in 14.5%. Cleanlist, which ran the study, is 17th at 9.1%. Across 280 matched answer pairs, 32% of the tools an engine named on August 3, 2026 were gone from its answer to the same question on August 7, 2026.

Last updated: August 7, 2026. Wave 1 was fielded August 3, 2026. Wave 2 was fielded August 7, 2026. This page was corrected on August 7, 2026 after a detection bug was found in the first version. Both raw datasets have been regenerated and republished under CC BY 4.0.

Correction, August 7, 2026: a detection bug inflated Cleanlist's own numbers

The first version of this page reported Cleanlist at 17.3% share of voice and rank 10 of 53 tools, including 41.7% and 45.8% share on Google AI Mode. Those figures were wrong. The corrected figures are 9.1% share, rank 17 of 51 tools, and 5.6% then 2.8% on Google AI Mode.

The bug. Google AI Mode returns its answer as markdown with source URLs written inline. The collector fed that raw markdown, link targets included, into the text it scanned for tool names. The pattern used to detect Cleanlist matches inside https://www.cleanlist.ai/blog/..., so a Google AI Mode answer that cited a cleanlist.ai page was scored as an answer that named Cleanlist. One of those answers recommends Clay in its visible text and contains the string "cleanlist" only inside a link target.

How it was found. 41.7% share on one engine against 6.9% on another, from the same questions on the same two days, is a gap worth opening the answers over before publishing. Reading the Google AI Mode answers behind that number showed the mentions were link targets.

Who it affected. Any brand whose own domain gets cited was inflated, so SyncGTM, Instantly and Amplemarket also came down. Cleanlist came down hardest, because cleanlist.ai is the most-cited vendor domain in the study. In share-of-voice points the drop was 8.2 for Cleanlist, 7.1 for SyncGTM and 4.2 for Amplemarket, against 1.6 for Apollo.

The fix. Detection now strips link targets and bare URLs from the answer text before scanning for brand names, and keeps them for the citation map. Both datasets were regenerated from the archived answers and republished. The exact patch is printed further down this page, and the corrected raw rows are linked at the bottom so the whole table can be recomputed without us.

The error made the company running the study look better than the data supports. That is why it is on this page rather than in a changelog.

Cleanlist ran this study, and Cleanlist is in it at #17

Cleanlist is a B2B data enrichment company. We designed the question set, ran the collection, wrote the detection code and are publishing the result. Cleanlist appears in this data at rank 17 of 51, with 9.1% share of voice, and 9.1% is down from what we first published. That is a conflict of interest and you should read the numbers with it in mind.

Two things reduce how much you have to take on trust. First, the full answer-level dataset is published, so every share figure on this page can be recomputed from the raw rows. Second, the numbers that make Cleanlist look bad are on this page in the same size type as the ones that do not: 3.3% on the largest buying job in the set, 0% on buying and pricing questions in wave 2, and 2.8% on Google AI Mode.

Method in one box

What was asked: 72 questions a real B2B buyer types, spread evenly across 6 buying jobs, 12 questions each.

Who was asked: ChatGPT (gpt-5.5), Claude (claude-sonnet-4-6), Gemini (gemini-3.5-flash) and Perplexity (sonar-pro), each called with web search enabled, plus Google AI Mode captured as a live SERP feature. United States location, English language, desktop.

When: wave 1 on 2026-08-03, wave 2 on 2026-08-07.

How much: 720 calls, 712 answers returned and archived (354 in wave 1, 358 in wave 2), 352 matched pairs, 51 distinct tools detected.

How tools were detected: a fixed dictionary of case-insensitive, word-boundary regular expressions applied to the answer text with link targets and bare URLs stripped out first. A tool counts once per answer regardless of how many times it is mentioned.

What this does not measure: a detection counts a mention, not a recommendation and not a rank. All answers were collected through the DataForSEO AI Optimization API and archived.

Which B2B data tools do AI engines recommend most in 2026?

Apollo is the most-named B2B data tool in AI answers, appearing in 75.6% of the 712 answers in this study, ahead of ZoomInfo at 57.9% and Clay at 46.5%. The table below is the share of all 712 answers that named each tool, combined across both waves, with the wave-by-wave split so you can see which brands moved. Share of voice here means the share of answers that named the tool at least once. It measures how often an answer engine says a name. It is not a ranking of the tools themselves, and it says nothing about whether a tool is any good.

#ToolAnswers naming itShareWave 1Wave 2Change
1Apollo53875.6%76.6%74.6%-2 pts
2ZoomInfo41257.9%59.3%56.4%-2.9 pts
3Clay33146.5%47.2%45.8%-1.4 pts
4Cognism27538.6%39.5%37.7%-1.8 pts
5Lusha25736.1%36.7%35.5%-1.2 pts
6HubSpot23432.9%38.1%27.7%-10.4 pts
7Hunter.io21129.6%32.2%27.1%-5.1 pts
8Salesforce16623.3%26.3%20.4%-5.9 pts
9Clearbit14520.4%22.9%17.9%-5 pts
10NeverBounce10314.5%14.1%14.8%+0.7 pts
11ZeroBounce8011.2%10.7%11.7%+1 pts
12Prospeo7610.7%11.6%9.8%-1.8 pts
13People Data Labs7610.7%10.7%10.6%-0.1 pts
14UpLead709.8%9.9%9.8%-0.1 pts
15Findymail679.4%9.6%9.2%-0.4 pts
16Smartlead669.3%9.9%8.7%-1.2 pts
17Cleanlist659.1%9.9%8.4%-1.5 pts
18Kaspr649%9.3%8.7%-0.6 pts
19Snov.io628.7%9.9%7.5%-2.4 pts
20Amplemarket598.3%10.5%6.1%-4.4 pts

The shape of that distribution matters more than any single row. Apollo is named in three quarters of all answers, the tool in 20th place is named in roughly one answer in twelve, and only nine of the 51 detected tools clear 20% share. The other 42 sit below it.

HubSpot posted the largest wave-to-wave move of any tool in the top 20, falling 10.4 points in four days on an unchanged question set. All 51 tools, broken out by engine, job and wave, are in the tool-share CSV linked at the bottom of this page.

How often do AI tool recommendations change?

They change a lot, and fast. Across 280 matched answer pairs on the four engines whose answers kept the same shape between waves, 32% of the tools named on August 3, 2026 were gone from the answer to the same question on August 7, 2026, and 30.4% of the August 7 names were not there four days earlier. The mean Jaccard similarity of the named-tool set between the two waves was 0.554, meaning a little over half of the union of the two answers was common to both.

Nothing about the questions changed. The prompt text, the model, the location, the language and the output cap were identical in both waves. The only variable was four days.

The number in plain language

Ask an AI engine the same B2B buying question twice, four days apart, and roughly one in three of the tools it named the first time will not be in the second answer.

Why was Perplexity excluded from the volatility headline?

Perplexity was excluded because its answers got dramatically shorter between the two waves, which makes a dropped tool impossible to distinguish from a shorter answer. Perplexity's mean answer length fell from 4,614 characters in wave 1 to 1,649 characters in wave 2, and its mean tools named per answer fell from 6.17 to 3.15. If an engine writes a third as much, it will name fewer brands whether or not its underlying view of the category changed. Counting that as churn would inflate the headline.

The exclusion rule was fixed in code before the result was read: any engine whose mean answer length moved by more than 25% between waves is dropped from the headline. Perplexity was the only engine that tripped it. The table below is the evidence, published in full so you can check the reasoning rather than take it.

EngineWaveAnswersMean tools namedMean answer charsMean sources cited
ChatGPTw1717.4241022.8
ChatGPTw2726.6342592.6
Claudew1674.8436916.2
Claudew2725.1136746
Geminiw1727.29452910.6
Geminiw2707.4458210
Google AI Modew1724.78275920.8
Google AI Modew2724.81264820.6
Perplexityw1726.17461419.1
Perplexityw2723.15164919.2

For completeness, the all-engine figures including Perplexity are: 352 matched pairs, mean Jaccard 0.529, 37.3% of wave-1 names gone, 28.8% of wave-2 names new. Those are the numbers a study trying to look impressive would have led with. They are not the headline here because the Perplexity component of them is contaminated.

That table also holds the most useful single row in the study. Google AI Mode cites more sources per answer than any other engine, 20.8 in wave 1 and 20.6 in wave 2, while naming fewer tools per answer than any other engine, 4.78 and 4.81. It reads the widest and says the least. That gap is the subject of the ghost citation section below.

Which AI engine changes its answer the most?

ChatGPT churns most among the shape-stable engines, dropping 37.8% of the tools it named in wave 1 within four days, and Claude churns least at 23.1%. Perplexity's 57.7% is the highest number in the table but it carries the answer-length problem described above, so treat it as unusable rather than as a finding.

EngineMatched pairsMean JaccardRetainedDroppedAdded% of wave-1 dropped
ChatGPT710.51532819913437.8%
Claude670.668249759123.1%
Gemini700.51935215716630.8%
Google AI Mode720.5223011411633.1%
Perplexity720.4341882563957.7%

Answer length does not explain the spread. Claude and Google AI Mode name almost the same number of tools per answer, 4.84 and 5.11 against 4.78 and 4.81, and they sit ten points apart on churn at 23.1% and 33.1%. Stability appears to be a property of the engine rather than of how much it writes.

Do different AI engines recommend different tools for the same question?

Yes, and the gaps are large enough that engine choice changes the shortlist. The same 72 questions produce very different name sets depending on who is answering. Clay is named in 69% of Gemini answers and 19.4% of Perplexity answers. Hunter.io is named in 49.3% of Gemini answers and 17.4% of Google AI Mode answers. People Data Labs is named in 21.8% of Gemini answers and 1.4% of Claude answers. Cleanlist is named in 15.3% of Perplexity answers and 4.2% of Google AI Mode answers.

ToolChatGPTClaudeGeminiGoogle AI ModePerplexity
Apollo80.4%73.4%84.5%79.2%60.4%
ZoomInfo69.9%51.8%56.3%60.4%50.7%
Clay52.4%38.1%69%53.5%19.4%
Cognism48.3%41%41.5%30.6%31.9%
Lusha45.5%30.9%43.7%33.3%27.1%
HubSpot43.4%28.8%44.4%22.2%25.7%
Hunter.io30.8%30.2%49.3%17.4%20.8%
Salesforce34.3%20.9%31.7%13.2%16.7%
Clearbit37.1%13.7%18.3%14.6%18.1%
NeverBounce23.1%5%27.5%13.2%3.5%
ZeroBounce22.4%5%16.9%8.3%3.5%
Prospeo14.7%7.9%16.2%2.1%12.5%
People Data Labs19.6%1.4%21.8%7.6%2.8%
UpLead15.4%6.5%11.3%10.4%5.6%
Findymail10.5%4.3%21.1%6.9%4.2%
Smartlead16.1%2.9%16.2%5.6%5.6%
Cleanlist7%12.9%6.3%4.2%15.3%
Kaspr8.4%4.3%13.4%10.4%8.3%
Snov.io11.9%12.9%4.9%4.9%9%
Amplemarket4.2%5.8%12%7.6%11.8%
RocketReach18.2%8.6%2.1%4.9%5.6%
SyncGTM2.8%3.6%15.5%7.6%9.7%
Seamless.AI11.9%10.8%9.2%0%6.9%
FullEnrich3.5%7.2%7.7%9.7%7.6%
MillionVerifier11.2%0.7%15.5%4.9%0.7%

Two rows show the mechanism. Seamless.AI is named in 11.9% of ChatGPT answers and 0% of Google AI Mode answers, the only zero in the table. MillionVerifier is named in 15.5% of Gemini answers and 0.7% of both Claude and Perplexity answers. A gap between zero and better than one answer in ten changes the shortlist a buyer sees.

Only Apollo and ZoomInfo are named in more than half of answers on all five engines. Everything else is engine-dependent.

Which tools do AI engines name for finding a contact's email or phone number?

For the 118 answers about finding contact data, Apollo is named in 87.3% and Lusha is second at 67.8%, the highest position Lusha reaches on any job. Full top-12: Apollo 87.3% · Lusha 67.8% · ZoomInfo 60.2% · Cognism 57.6% · Hunter.io 48.3% · Clay 30.5% · Kaspr 29.7% · RocketReach 19.5% · Snov.io 19.5% · NeverBounce 14.4% · UpLead 14.4% · ZeroBounce 13.6%.

This job has the most distinctive shape of the six. Kaspr appears in its top twelve and in no other job's, and Clay drops from third on enrichment at 66.7% to sixth here at 30.5%. Engines appear to separate "find one person's contact details" from "process a list", and they name different vendors for each. Cleanlist is at 8.3% then 8.6% here and does not reach the top twelve.

Which tools do AI engines name for building a prospect list?

For the 120 answers about building a prospect list, Apollo is named in 92.5%, the highest concentration for any tool on any job in the study. Full top-12: Apollo 92.5% · ZoomInfo 75.8% · Clay 55.8% · HubSpot 42.5% · Lusha 39.2% · Hunter.io 38.3% · Cognism 37.5% · Salesforce 30.8% · NeverBounce 24.2% · Smartlead 23.3% · Clearbit 20% · Prospeo 18.3%.

This is the most concentrated job in the set. The top two names each appear in more than three quarters of answers, which means a buyer asking an AI engine how to build a list gets a near-identical opening pair no matter how they phrase it. Cleanlist is at 3.3% in both waves here, its weakest job-level result and the joint-weakest result of its six job scores, on the job that has the most buyers in it.

Which tools do AI engines name for enriching and cleaning data?

For the 120 answers about enriching and cleaning data, Apollo leads at 85.8%, followed by ZoomInfo at 74.2%, with Clay and HubSpot tied at 66.7%. Full top-12: Apollo 85.8% · ZoomInfo 74.2% · Clay 66.7% · HubSpot 66.7% · Clearbit 61.7% · Cognism 60% · Salesforce 38.3% · Lusha 37.5% · People Data Labs 37.5% · Hunter.io 29.2% · Cleanlist 24.2% · FullEnrich 19.2%.

Enrichment is the job where HubSpot ranks highest of the six. At 66.7% it ties Clay for third and beats every specialist data vendor except Apollo and ZoomInfo, and Salesforce at 38.3% comes in ahead of People Data Labs and Hunter.io. Engines appear to treat "clean my CRM data" and "enrich a CSV" as adjacent problems and answer both with the CRM. This is also Cleanlist's strongest job, at 26.7% in wave 1 and 21.7% in wave 2, and it still fell five points in four days.

For the 120 answers about AI agents and natural-language prospecting, Apollo is named in 61.7% and Clay in 54.2%, the closest first-and-second gap of any job at 7.5 points. Full top-12: Apollo 61.7% · Clay 54.2% · HubSpot 46.7% · Salesforce 40% · ZoomInfo 35.8% · Amplemarket 20.8% · SyncGTM 16.7% · 11x 15% · Clearbit 14.2% · Hunter.io 13.3% · Cognism 13.3% · Crustdata 13.3%.

Two observations. First, the incumbents hold this job too: four of the five names that lead the list-building job also lead the agent job, which cuts against the idea that AI-native positioning automatically wins AI answers. Second, this job carries names that appear nowhere else, with 11x and Crustdata in its top twelve and in no other job's. Cleanlist sits at 1.7% in wave 1 and 5% in wave 2 on the job closest to what we build, outside the top twelve in both waves, which is a result worth stating plainly rather than explaining away.

Which tools do AI engines name for email deliverability and data quality?

For the 119 answers about data quality, accuracy and deliverability, Apollo leads at 43.7%, its lowest score on any job, and two email verification vendors reach the top five. Full top-12: Apollo 43.7% · ZoomInfo 30.3% · Clay 30.3% · NeverBounce 29.4% · ZeroBounce 29.4% · Hunter.io 24.4% · Cognism 23.5% · Findymail 17.6% · Prospeo 15.1% · MillionVerifier 14.3% · Cleanlist 13.4% · Lusha 12.6%.

This is the most open job in the study. Apollo's 43.7% is the lowest first-place score of any of the six jobs, more than 40 points below its 92.5% on list building, and three email verification specialists appear in the top twelve. Questions about bounce rates and accuracy pull engines away from the large platforms and toward vendors that do one thing. Cleanlist is at 15.3% then 11.7% here, its second-best job.

Which tools do AI engines name for pricing, buying and switching questions?

For the 115 answers about pricing, buying and switching, Apollo is named in 82.6% and ZoomInfo in 71.3%, even though several of the questions asked explicitly for cheaper options or alternatives. Full top-12: Apollo 82.6% · ZoomInfo 71.3% · Lusha 53% · Clay 40.9% · Cognism 40% · Hunter.io 24.3% · HubSpot 21.7% · UpLead 20.9% · Salesforce 14.8% · Prospeo 12.2% · RocketReach 11.3% · Snov.io 11.3%.

That is the counterintuitive result in the job tables. Asking an engine for a cheaper option than the market leader still surfaces the market leader in most answers, usually as the baseline the alternatives are measured against. A vendor competing on price is therefore competing inside a passage that names the incumbent first. Cleanlist scored 3.6% in wave 1 and 0% in wave 2 on this job, its only zero anywhere in the study.

Which tools kept their place between the two waves?

Apollo held 89.2% of its wave-1 appearances into wave 2, the highest retention of any tool in the table below, and Lemlist held 25.9%, the lowest among tools with at least 20 wave-1 appearances. Retention here is the share of the answers that named a tool in wave 1 where the same engine still named it in wave 2 for the same question.

ToolWave-1 appearancesRetainedRetention
Apollo26924089.2%
ZoomInfo20917382.8%
Clay16612474.7%
Cognism13811079.7%
HubSpot1347052.2%
Lusha1299472.9%
Hunter.io1127567%
Salesforce924548.9%
Clearbit814353.1%
NeverBounce493469.4%
Prospeo411843.9%
ZeroBounce382463.2%
People Data Labs372259.5%
Amplemarket371540.5%
Snov.io351542.9%
Cleanlist351748.6%
Smartlead351440%
UpLead341647.1%
Findymail341955.9%
Kaspr331236.4%
Seamless.AI321340.6%
SyncGTM311445.2%
FullEnrich281242.9%
Lemlist27725.9%
RocketReach261453.8%
MillionVerifier221045.5%
LeadIQ21733.3%
Anymailfinder191157.9%
Saleshandy19631.6%
Dropcontact18844.4%

Retention tracks size, but not cleanly. HubSpot had 134 wave-1 appearances and kept 52.2% of them, worse than NeverBounce, which had 49 and kept 69.4%. Salesforce had 92 and kept 48.9%, worse than Anymailfinder, which had 19 and kept 57.9%. Being named often is not the same as being named durably.

Cleanlist kept 17 of its 35 wave-1 appearances, 48.6%, which puts it just below Salesforce and sixteenth of the thirty tools in this table.

Which websites do AI engines cite when they answer B2B data questions?

Excluding Google's own grounding redirect infrastructure, Reddit is the most-cited domain in this study with 279 citations across ChatGPT, Google AI Mode and Perplexity, followed by YouTube at 252. The first B2B data tool vendor's own domain in the list is cleanlist.ai, at 213 citations across four of the five engines.

#DomainCitationsEngines citing it
1vertexaisearch.cloud.google.com1463Gemini
2www.reddit.com279ChatGPT, Google AI Mode, Perplexity
3www.youtube.com252Google AI Mode, Perplexity, ChatGPT
4google.com240Google AI Mode
5www.cleanlist.ai213ChatGPT, Claude, Google AI Mode, Perplexity
6pipeline.zoominfo.com212Claude, Google AI Mode, Perplexity
7syncgtm.com158ChatGPT, Claude, Google AI Mode, Perplexity
8www.apollo.io131ChatGPT, Claude, Google AI Mode, Perplexity
9www.cognism.com120ChatGPT, Claude, Google AI Mode, Perplexity
10www.linkedin.com114ChatGPT, Google AI Mode, Perplexity
11www.amplemarket.com98ChatGPT, Claude, Google AI Mode, Perplexity
12prospeo.io96ChatGPT, Claude, Perplexity
13www.clay.com95ChatGPT, Claude, Google AI Mode, Perplexity
14salesmotion.io79Claude, Google AI Mode, Perplexity
15www.saleshandy.com78Claude, Google AI Mode, Perplexity
16www.salesforge.ai77Claude, Google AI Mode, Perplexity
17www.unifygtm.com73ChatGPT, Claude, Google AI Mode, Perplexity
18www.autobound.ai72Claude, Google AI Mode, Perplexity
19www.landbase.com70Claude, Google AI Mode, Perplexity
20instantly.ai65Google AI Mode, Perplexity

The top entry is infrastructure, not a publisher. vertexaisearch.cloud.google.com is the grounding redirect Gemini returns instead of a destination URL, so the underlying sources behind those 1,463 citations are not visible in the response. That is a real limit on what any study can say about Gemini's sourcing, and it is the reason Gemini is absent from the "engines citing it" column for every domain below it.

Two structural points survive that limit. Vendor-owned marketing sites dominate the list, with fifteen of the twenty domains above belonging to a tool vendor. And community content still out-cites all of them individually, since Reddit's 279 citations are higher than any single vendor domain in the table.

Why is Cleanlist cited 92 times by Google AI Mode and named in 6 of its answers?

This is the finding the correction produced, and it is the most useful thing in the study. The citation map was built from the URL list rather than the answer prose, so it was never touched by the detection bug. Once the naming figures were fixed, the two measurements separated, and the gap between them is large.

Google AI Mode cited cleanlist.ai 92 times across the two waves. Google AI Mode wrote the word Cleanlist in 6 of its 144 answers. The engine reads the source constantly and almost never says the name.

EngineCitations of cleanlist.aiAnswers naming CleanlistAnswers
Google AI Mode926144
Perplexity7822144
Claude3318139
ChatGPT1010143
Gemininot visible9142

Gemini's column is blank because all 1,463 of its citations are grounding redirects, so its actual sources cannot be read.

The mechanism is visible in the answer-shape table above. Google AI Mode cites 20.8 sources per answer in wave 1 and 20.6 in wave 2, more than any other engine, while naming 4.78 and 4.81 tools per answer, fewer than any other engine. It reads about twenty pages and names about five products. Most of what it reads is discarded before the visible answer is written, and a vendor page can easily be in the read set and out of the named set.

Perplexity is the counter-case at the other end. It cited cleanlist.ai 78 times, close to Google AI Mode's 92, and named Cleanlist in 22 answers against Google AI Mode's 6. ChatGPT cited cleanlist.ai only 10 times and named Cleanlist in 10 answers. The ratio of citations to names is not a constant across engines, which means it is a property worth measuring per engine rather than assuming.

What does a ghost citation mean for a vendor?

Being the source an engine reads is not the same as being the answer it gives. Those two outcomes are produced by different work, and a vendor that measures only one of them will draw the wrong conclusion.

A study that counted only citations would have reported that Cleanlist is doing well on Google AI Mode, since cleanlist.ai is the fifth most-cited domain in the study and the most-cited vendor domain. A study that counted only names would have reported that Google AI Mode ignores Cleanlist, at 4.2% share against 15.3% on Perplexity. Both readings are available from this dataset and both are incomplete. The pair is the finding.

The practical split, stated as a hypothesis this data supports rather than a law it proves:

  • Getting cited is a retrieval outcome. It responds to having a page that matches the question, is crawlable, and is specific enough to be worth pulling into a grounding set. cleanlist.ai earned its 213 citations from 12 distinct URLs, all twelve of them long-form comparison, pricing or how-to guides on the blog.
  • Getting named is a generation outcome. The engine has already retrieved twenty sources and is choosing five product names to write down. What survives that step is the brand the sources themselves talk about, not the brand that wrote the source.

The uncomfortable version for anyone doing answer engine optimization: publishing the page that an engine cites can result in the engine naming your competitors, using your page as the evidence. One of the Google AI Mode answers in this study cites cleanlist.ai and recommends Clay in its visible text. That answer is in the published dataset.

This is one vendor, in one category, on five engines, over four days. It is not a general law of AI search, and the direction of causality is not established here. It is a measured gap large enough that other vendors should check whether they have the same one, and the method for checking is published below.

Where does Cleanlist rank in its own study?

Cleanlist ranks 17th of 51 tools with 9.1% share of voice across both waves, and it fell from 9.9% in wave 1 to 8.4% in wave 2. The distribution behind that number is lopsided. Cleanlist is named in 15.3% of Perplexity answers and 4.2% of Google AI Mode answers.

By wave and engine: Perplexity 18.1% then 12.5%, Claude 11.9% then 13.9%, ChatGPT 7% then 6.9%, Gemini 6.9% then 5.7%, Google AI Mode 5.6% then 2.8%. Claude is the only engine where Cleanlist rose between the waves.

By job, Cleanlist's best result is enrichment at 26.7% then 21.7%. Its worst is building a prospect list at 3.3% in both waves, which is the job with the largest buyer population in the set, and buying and pricing at 3.6% then 0%. Contact finding is 8.3% then 8.6%, data quality 15.3% then 11.7%, and AI agents 1.7% then 5%. Cleanlist reaches the top twelve on two of the six jobs.

Cleanlist retained 48.6% of its wave-1 appearances, below Apollo (89.2%), ZoomInfo (82.8%), Cognism (79.7%), Clay (74.7%), Lusha (72.9%), NeverBounce (69.4%), Hunter.io (67%), ZeroBounce (63.2%), People Data Labs (59.5%), Anymailfinder (57.9%), Findymail (55.9%), RocketReach (53.8%), Clearbit (53.1%), HubSpot (52.2%) and Salesforce (48.9%).

The first version of this page reported 17.3% and rank 10, and reported Google AI Mode as Cleanlist's strongest engine at 41.7% then 45.8%. Google AI Mode is in fact Cleanlist's weakest engine by naming, at 4.2%, and its strongest by citation, at 92 citations. The entire "AI Mode stronghold" in the first version was the detection artifact.

How was this study measured?

Each of the 72 prompts was sent to five engines in each wave, 720 calls in total, of which 712 returned a parseable answer. The returned answer text was scanned for tool names using a fixed dictionary of case-insensitive, word-boundary regular expressions, applied after link targets and bare URLs were stripped out. ChatGPT, Claude, Gemini and Perplexity were called through their respective live LLM endpoints with web_search enabled and an output cap of 1,500 tokens. Google AI Mode was captured through a live SERP endpoint rather than a chat API, because it is a search feature and not a model with a public chat interface. All five used a United States location and English language; the SERP capture used desktop.

The 72 questions were split 12 apiece across six buying jobs: finding contact data, building a prospect list, enriching and cleaning data, AI agents and natural-language search, data quality and deliverability, and buying, pricing and switching. Sample prompts include "How do I build a cold outreach list from scratch in 2026?", "What is waterfall enrichment and which tools do it best?", "How do I stop my cold emails from bouncing?" and "Is there an MCP server for B2B lead data and enrichment?".

Citations were collected separately from names. For the four chat engines, citation URLs come from the annotation objects attached to each answer section. For Google AI Mode, they come from the link targets in the returned markdown. The citation map was therefore never affected by the detection bug described in the next section.

What was the detection bug, and how is it fixed?

The first version of this page ran brand detection over the raw markdown that Google AI Mode returns, with inline source URLs still in it. A brand pattern like the one for Cleanlist matches inside a URL, so a citation of https://www.cleanlist.ai/blog/2026-03-31-best-data-enrichment-tools-2026 scored as an answer naming Cleanlist. Cleanlist was the biggest beneficiary because cleanlist.ai is the most-cited vendor domain in the study, but any brand whose domain appears in a citation was affected, including SyncGTM, Instantly and Amplemarket.

The fix strips link targets and bare URLs from the text before brand detection, while keeping them for the citation map:

// Google AI Mode returns its answer as markdown with source URLs inline. Those URLs must
// NOT reach brand detection: a pattern like /\bcleanlist\b/i matches inside
// "https://www.cleanlist.ai/blog/...", so a CITATION would be scored as the engine NAMING
// the tool. Strip link targets and bare URLs from the text used for detection; keep them
// for the citation map.
const stripUrls = (md) =>
  md.replace(/\]\((https?:\/\/[^)]*)\)/g, ']()').replace(/https?:\/\/\S+/g, '');

Both published datasets were regenerated from the archived answers with this fix in place. No answers were re-collected, so the correction changes only how the same archived text was counted.

If you are running your own AI visibility measurement, this is the specific thing to check first. Any engine that returns markdown rather than plain text plus a separate source list will put brand domains into the character stream you are scanning. The bias is not random. It systematically favors whichever brand publishes the pages that the engine cites, which is disproportionately likely to be the brand running the measurement.

What are the limits of this study?

The largest limit is that a detection counts a mention, not a recommendation and not a rank. If an engine writes "Apollo is cheap but its phone data is thin, so use X instead", both Apollo and X score a mention. Any claim on this page about what engines "recommend" should be read as what they name.

Five more limits, stated so you can discount the results correctly:

  • URL text in the answer body. Detection now strips link targets and bare URLs, because without that step a citation of a vendor's domain is scored as the engine naming that vendor. This study's first published version had exactly that bug and it inflated the publisher's own share by 8.2 points. Any AI visibility measurement that regex-scans answer text should be assumed to carry this bias until it shows the stripping step.
  • Dictionary blindness. Detection uses a fixed list of tool-name patterns. A tool that is not in the dictionary is invisible to this study no matter how often engines name it. 51 tools were detected; the dictionary is not the whole market.
  • One locale. Every query used a United States location and English language. Results for a UK, German or Indian buyer are not measured here and should not be assumed to match.
  • Two waves is a snapshot. Four days apart is enough to demonstrate instability. It is not enough to establish a trend, a direction, or a seasonal pattern. A tool that fell between these two waves may be flat over a month.
  • Model versions move. The four chat engines were pinned to specific model versions. Providers update models without notice, so a replication run weeks later is testing a partly different system.

Can I reproduce this study?

Yes. Both datasets are published under CC BY 4.0, which means you can redistribute, remix and build on them commercially as long as you credit Cleanlist. Both files below are the corrected versions, regenerated after the detection fix.

  • Answer-level dataset (712 rows): one row per AI answer, with columns wave, wave_date, engine, question_id, job, question_slug, tools_named, n_tools, n_sources, cleanlist_named.
  • Tool-share dataset (2,307 rows): one row per tool by engine by job by wave, with answers_naming_tool, answers_total and share_pct. Rows where engine or job is ALL are the aggregate cuts.

Three notes for anyone recomputing. The job column uses the internal codes find, build, Enrich and clean data, ai, data and Buying, pricing and switching, which map to finding contact data, building a prospect list, enriching and cleaning data, AI agents and natural language, data quality and deliverability, and buying, pricing and switching. The question_slug is a truncated 44-character slug of the prompt, not the full prompt text, so joining on it is safe within this dataset but not sufficient to re-send the exact query. And the cleanlist_named column is the corrected one: it sums to 65, not the 122 the first version of this page was built on. For the full prompt set and the analysis script, including the stripUrls patch above in context, email victor@cleanlist.ai and we will send both.

If you replicate this and get a different answer, publish it. A category with two independent measurements is better off than a category with one, including when the second one contradicts the first.

What should a buyer do with this?

Do not treat a single AI answer as a shortlist. If a third of the tools an engine names today are gone from the same answer in four days, then the specific set you were shown is partly an artifact of when you asked. The stable signal in this data sits at the top of the list. Apollo, ZoomInfo, Cognism and Clay each held 74.7% or better of their wave-1 appearances, while every other tool in the published retention table held somewhere between 25.9% and 72.9%.

Three practical moves. Ask two engines, not one, because Clay's 69% on Gemini and 19.4% on Perplexity are the same tool on the same questions. Ask the same engine twice a few days apart and take the intersection. And treat everything below the top four as a candidate to verify rather than a ranking to trust, because that is the band where the churn lives.

What should a vendor do with this?

Stop treating one good AI answer as a position. In this study the mean Jaccard similarity between two runs of the same question was 0.554, so a little over half of the combined tool set across the two runs was common to both and the rest was churn. A screenshot of an engine naming your brand is a sample of one from a distribution that moves every few days, and it is not evidence that anything changed.

Measure citations and names as two separate numbers. Cleanlist's own data is the argument: 92 citations from Google AI Mode against 6 answers that write the name, versus 10 citations from ChatGPT against 10 answers that write the name. A single blended "AI visibility" score would have averaged those into something meaningless. Reported separately, they point at different work: one page-level retrieval problem and one brand-in-the-corpus problem.

And check your own detection for the URL bias before you trust your dashboard. If your measurement scans answer text for your brand name and your brand name is in your domain, some share of what you are calling visibility is your own citations counted twice.

Enrich 30 leads free

Upload a CSV, run it through the Cleanlist waterfall, and keep the results. Free plan is 30 credits a month, no credit card. Search costs 0 credits.

Enrich 30 leads free

FAQ

What is an AI visibility index?

An AI visibility index measures how often AI answer engines name a given brand when real buyers ask category questions. Cleanlist's index for B2B data tools works by putting a fixed set of 72 buying questions to five engines, capturing every answer, and counting the share of answers that name each tool. Share of voice is the headline metric: Apollo 75.6%, ZoomInfo 57.9%, Clay 46.5% across 712 answers collected on August 3 and August 7, 2026. The index measures naming frequency in generated answers. It does not measure product quality, market share, revenue or customer satisfaction, and a mention inside a critical sentence counts the same as a mention inside a recommendation.

How many AI answers were in this study?

712 answers: 354 collected on August 3, 2026 and 358 collected on August 7, 2026, across ChatGPT, Claude, Gemini, Google AI Mode and Perplexity. Of those, 352 form matched pairs, meaning the same engine answered the same question in both waves. The volatility headline uses a 280-pair subset that excludes Perplexity, because Perplexity's mean answer length fell from 4,614 to 1,649 characters between waves and a dropped tool cannot be separated from a shorter answer. 51 distinct tools were detected across the full set. Every answer was archived and the answer-level dataset is published under CC BY 4.0.

Which AI engine is most stable for B2B tool recommendations?

Claude is the most stable of the five engines in this study, dropping 23.1% of its wave-1 tool mentions four days later and posting the highest mean Jaccard similarity at 0.668. ChatGPT is the least stable of the shape-stable engines, dropping 37.8% of the tools it named in wave 1, with Google AI Mode at 33.1% and Gemini at 30.8%. Answer length does not explain the spread: Claude and Google AI Mode name almost the same number of tools per answer, 4.84 and 5.11 against 4.78 and 4.81, and still sit ten points apart on churn. Perplexity's 57.7% drop rate is excluded from this comparison because its answers got roughly two thirds shorter between waves.

Can an AI engine cite a website without naming the company that owns it?

Yes, and the gap can be large. In this study Google AI Mode cited cleanlist.ai 92 times across both waves and wrote the word Cleanlist in 6 of its 144 answers. Perplexity cited cleanlist.ai 78 times, a similar number, and named Cleanlist in 22 answers. The mechanism is visible in answer shape: Google AI Mode cites about 20 sources per answer and names about 5 tools, so most of what it retrieves never reaches the visible text. For a vendor this means citation share and naming share are two separate metrics with two separate causes, and one of the Google AI Mode answers in the published dataset cites cleanlist.ai while recommending Clay in its visible text.

Is Cleanlist's own ranking in this study reliable?

Treat it with the same scepticism you would apply to any vendor publishing about itself, and note that the first version of this page overstated it. Cleanlist designed the questions, ran the collection, wrote the detection code and published the result, and a bug in that detection code inflated Cleanlist's share from 9.1% to a published 17.3% before it was caught. Cleanlist appears in the corrected data at rank 17 of 51 with 9.1% share of voice. Three things make the figure checkable: the corrected answer-level dataset is public so the number can be recomputed independently, the question set was fixed across six jobs before any results were seen, and the weak results are published alongside the strong ones. Cleanlist scores 3.3% on prospect-list building, 0% on buying and pricing questions in wave 2, and 2.8% on Google AI Mode in wave 2.

No. Detection in this study is a case-insensitive, word-boundary regular expression run over the answer text, so a tool counts as named whether the engine praised it, criticised it, or listed it as the thing to avoid. An answer that says "Apollo has weak phone coverage in Europe, use something else" scores a mention for Apollo. This is the single most important limit on every number in the report, which is why the raw tools_named column is published per answer: anyone who wants a stricter definition, such as sentiment-scored or position-weighted mentions, can build it from the same rows rather than taking our count.

Try it now

Run this on your own list

Upload a CSV, enrich every row across 15+ providers, and export the result back to your CRM. First 30 rows free.

Enrich 30 rows free

30 credits free every month · No credit card

30 credits included. No credit card required. Set up in 5 minutes.