269 AI Answer Sentences Traced to Our Pages: 16 Name Us, 253 Name Somebody Else

We traced 269 sentences from 37 Google AI answers back to the Cleanlist pages they cited. 16 came from a sentence naming Cleanlist and 253 did not. Open dataset, CC BY 4.0.

Victor Paraschiv

Victor Paraschiv

Co-Founder & CMO

24 min read

On September 7, 2026, 37 Google AI answers cited a cleanlist.ai page. We took the text of those answers, split it into sentences, and matched each sentence back to the closest sentence on the page it cited. 269 sentences matched. 16 trace to a sentence containing the word Cleanlist and 253 do not. The most-reused thing we own is a line we wrote about a competitor. This is one night, one country, 21 pages, and we are the subject.

Last updated: September 7, 2026. Every answer in this study was collected on September 7, 2026.

Method, including the part that is an inference

What was collected. 240 B2B buyer queries on Google US desktop plus 60 Google AI Mode questions, DataForSEO, location_code 2840, English, desktop, one live pull per query, September 7, 2026. 37 answers cited a cleanlist.ai URL in the answer body or the reference list: 32 AI Overviews and 5 AI Mode answers, covering 21 distinct pages across 39 page-and-answer pairs.

How sentences were matched. Each cited page was read at git HEAD, not from the live site and not from the working tree, because six of the pages were edited later the same morning. MDX prose and the string contents of our TypeScript page-data files were segmented at headings, list items, table cells and sentence terminators, dropping segments under 38 characters. That leaves 4,483 eligible source sentences across the 21 pages, of which 697 contain the word Cleanlist. Answer bodies were segmented the same way after stripping citation markup.

The similarity rule. Each answer sentence is scored against every source sentence on the pages that answer cited, using Dice coefficient over lowercased word tokens of four characters or more with a stoplist. The best-scoring source sentence is kept when the score reaches 0.30. 269 answer sentences clear that bar, and every one is a row in the published dataset with its score.

This match is an inference, and it is labelled as one throughout. Google returns no provenance signal. Nothing in the API says "this sentence came from that sentence". High token overlap between an answer sentence and a sentence on a page the answer cited is evidence of reuse and not proof of it. Where this page says a sentence was lifted, read "matched by token overlap to". A control below shows how often that inference fires on a page the answer never cited.

Sample. 37 answers, 21 pages, one day, one locale, one device, observational. Two subsets do most of the work: 11 of the 39 page-citation pairs land on a single listicle and 3 more on a single alternatives page.

What did we measure?

Cleanlist has the opposite of the usual citation problem. Google reads us. In the 240-query buyer census run the same night, cleanlist.ai was the seventh most-cited domain in AI Overviews, cited on 32 of them, and Cleanlist was named in the answer text of 6 to 11 of roughly 230, depending on which of two identical pulls you read. Plenty of reading, almost no recommending.

The usual explanation is that we are not writing the sentence the engine wants. This study tests a different one. When the engine reads our page and writes an answer, whose sentences is it reusing, ours or the ones we wrote about other vendors?

So we traced it: answer sentence, best-matching source sentence, page, score. Then we counted one thing, which needs no model once the trace exists. How many of the source sentences contain the word Cleanlist?

We ran this, we are in it, and we lose

Cleanlist is a B2B data enrichment company. We chose the query set, ran the collection, wrote the matching code, and are publishing a result in which our own content donates 253 sentences to answers that then recommend other vendors. The dataset is under CC BY 4.0 so every count can be recomputed without us, including the ones that make us look worst.

Two companies need a disclosure. Lusha is a supplier Cleanlist buys data from, and it appears here only because we ranked it on our own listicles, sixth on one and fourth on another. Our own enrichment page already describes it as an upstream provider inside the Cleanlist waterfall. It is reported as an observation about our pages and never as a competitor. Cognism appears as a market participant in a census of our own sentences. Neither appears in any shortlist, and this study authors no shortlist.

How many of the sentences Google took from our pages name us?

16 of 269, which is 5.9%. The other 253 do not contain our name anywhere.

The ratio is the finding and the integers are the fragile part of it. Move the similarity floor and the counts move while the share barely does: 7.3% at Dice 0.20, 5.9% at 0.30, 3.8% at 0.40 and 5.5% at 0.50. Swap the metric entirely for character 4-gram Jaccard, which shares no arithmetic with the first, and the share lands between 2.4% and 4.6%. Whatever the matcher is doing, it is not the reason the number is small.

There is also a version with no matcher in it. Of the 37 answers that cited a cleanlist.ai page that night, 6 named Cleanlist in the answer prose and 31 did not. Across the 39 page-citation pairs the naming count is also 6. That is a string search on the answer text, and it agrees with the traced figure.

Per answer, the picture holds. At Dice 0.40, 30 of the 35 answers with any trace at all, which is 85.7%, contain no traced sentence whose source names us. At 0.30 it is 24 of 37 (64.9%). At 0.20, where the matcher is loosest, it is 20 of 37 (54.1%). The margin narrows as the floor drops, which is expected, and the direction never changes.

Could this just be the composition of our own pages?

This is the objection that mattered most, and it is the one place this study has a control.

If our pages rarely say "Cleanlist", an answer drawing sentences at random would rarely land on one that does, and there would be nothing to explain. So we counted the base rate. Across the 21 cited pages, sentences containing the word Cleanlist are 15.8% of the eligible sentence corpus (678 of 4,293 under the verification pass's segmentation, 697 of 4,483 under the published one). They are 3.8% to 7.3% of the sentences the answers actually match to.

Treating each traced sentence as a draw from the page's own sentence pool, the binomial lower-tail probability of seeing this few Cleanlist-naming sources is 3.0e-7 at Dice 0.20, 6.6e-7 at 0.30, 1.6e-5 at 0.40 and 0.019 at 0.50. Under a sharper provenance test, where the cited page has to beat every other Cleanlist page we know of before the match counts, it holds at p = 0.005 at Dice 0.40. The engine is not mirroring the composition of our pages. It selects against the parts of them that name us.

The honest limit on the matcher sits next to that result. We ran a random-page control, matching each citing answer against a Cleanlist page it did not cite. 61% of sentences clear the equivalent of Dice 0.20, 24.8% clear 0.30, 6.2% clear 0.40 and 1.1% clear 0.50. At the 0.30 floor used for the headline, roughly a quarter of matches would fire on a page the answer never read. That is why the depletion result, which has a base rate to compare against, is better evidence than the raw ratio, which does not.

Which sentence of ours is reused most, and whose is it?

Not ours. The sentences with the highest reuse across separate answers are, almost without exception, ones we wrote about other companies.

Source sentence on a Cleanlist pageDescribesAnswers
Contact data: Verified email addresses, direct dial phone numbers, mobile numbers, job titlesgeneric feature line5
Clearbit / HubSpot Breeze Intelligence, Best for HubSpot UsersClearbit4
Diamond Data: phone-verified mobile numbers with 98% connect rateCognism4
Cognism, Best for European Data and GDPR ComplianceCognism3
Cognism: Best for teams selling into EMEA that need phone-verified mobiles and GDPR workflows.Cognism3
Cognism has built its reputation on two things: phone-verified mobile numbers (Diamond Data) and strong European contact coverage.Cognism3
Apollo combines a 275M+ contact database with built-in email sequencing, which makes its $0 plan the most complete free sales toolApollo3
Apollo bundles lead enrichment with sequencing and a dialer from $49 per user per month on BasicApollo3
If you want to build custom enrichment logic with multiple providers, Clay gives you that flexibility.Clay3
Cleanlist is our pick as the best Clay alternative for most teams: the same multi-provider waterfall comes pre-built and verified, from $79/mo, with no separate data credits on top.Cleanlist2

Every row except the last is a sentence we wrote to be fair to a competitor, and the highest-reuse sentence that names us sits at two answers.

Single-sentence reuse counts are the least stable statistic on this page, and we would rather say so than let a reader find out. An earlier internal pass reported one Cognism line reused seven times. That number came from a segmentation bug: the splitter failed to break between an H3 heading and the bullet underneath it, glued them into one long pseudo-sentence, and produced a token set fat enough to be the best match for almost anything. Split properly, the bullet alone is the best match in 3 answers and the heading alone in 3. No single sentence in this corpus reaches 7 under the token rule, and the maximum is 5.

Two forms of the claim survive a change of metric, and those are the ones to quote. Under character 4-gram Jaccard, the single most-reused string in the corpus is our H3 "Cognism, Best for European Data and GDPR Compliance", the best match in 8 answers, ahead of our Clearbit heading at 5 and our ZoomInfo heading at 4. At the brand level under token Dice, 12 of the 37 citing answers draw their best-matching source sentence from something we wrote about Cognism, 11 from something we wrote about ZoomInfo and 8 from something we wrote about Clearbit, against 13 for Cleanlist.

The mechanism, stated as the interpretation it is: a vendor listicle that gives every competitor a compact, complete, correctly-shaped "Brand: Best for X. It does Y, from $Z" sentence has manufactured a drop-in answer bullet for each of those competitors. The engine arrives looking for that shape and finds a page full of them.

The model-free version needs no matcher. Across the 21 cited pages there are 9 "Brand: Best for" constructions with Cleanlist as the grammatical subject and 45 with a rival as the subject, a ratio of 1 to 5. On the single most-cited page, which carries 11 of the 39 citations, the split is 1 of ours against 18 rivals: 10 prose bullets and 8 H3 headings, each a complete, liftable answer line for a company that is not us.

No, and the correlation runs the wrong way for the intuitive fix, so the intuitive fix is wrong.

The obvious reading of everything above is "we mention competitors too much, cut the mentions". The data does not support that. Across the nine pages with enough citations to rank, the Spearman correlation between rival-mention density and our own naming rate is +0.18 on the corrected collection and +0.31 on the superseded one. Positive. Pages where rivals are named more often are, if anything, marginally more likely to get us named.

Look at the two ends. The page named in 3 of 3 answers that cited it, /alternatives/clay, is wall-to-wall competitor: 31 Cleanlist mentions against 201 rival brand mentions. The page with the lowest rival density in the set, /glossary/email-append at 0.42 rival mentions per Cleanlist mention, was named in 0 of 2.

Both directions need the same caveat. Nine points cannot settle a correlation, +0.18 is weak, and the interval spans zero comfortably. What the data supports is a negative: no evidence that cutting competitor mentions gets you named, and none that adding them does either. Deleting rival coverage from a comparison page costs you the reason the page ranks and buys nothing this study can measure.

What actually predicts whether a page gets its owner named?

The safest answer is that we do not know, and here is why we are not printing a number.

The internal diagnostic reported a Spearman of +0.68 for "a shaped sentence exists, early". It does not survive independent operationalization. Recomputed against the diagnostic's own published ordinal coding, it is +0.81 on the superseded collection and +0.64 on the corrected one, so the published value depends on which pull you read. Coded mechanically from the page text instead of by hand, the same construct flips sign: -0.12 for "shaped sentence present and early", -0.16 for the count of our shaped sentences, -0.30 for how early the first one appears. Five of the nine most-cited pages carried a Cleanlist-subject "Best for" construction at HEAD on the night of the measurement, and three of those five were named zero times. Publishing a coefficient here would be publishing the coding choice.

Two rules that looked solid also fail against this corpus.

The earliness rule fails. The one shaped bullet Google demonstrably re-emitted, in form and with our real price, sits at character 28,829 of the enrichment listicle. Depth did not stop it being used.

The literal string rule fails. The page named 3 of 3, /alternatives/clay, contains no "Cleanlist: Best for" sentence anywhere. What it has is a self-contained summary at character 66 with Cleanlist as the grammatical subject, a segment and a real price: "Cleanlist is our pick as the best Clay alternative for most teams: the same multi-provider waterfall comes pre-built and verified, from $79/mo, with no separate data credits on top." The engine built the "Best for" shape out of that on its own.

What the evidence supports is narrower and less quotable. When Google names us it emits our shape back, and that shape is a self-contained sentence with our name as subject, a narrow segment and a number. Verbatim from the corrected pull, on "clay alternatives" and "clay alternative":

Cleanlist: Best for pre-built waterfall data enrichment without paying separate extra data credits.

And on "lead enrichment tools":

Cleanlist: Best for raw data accuracy, using a waterfall approach that yields high verified email and phone match rates starting at $79/month.

The second carries our real price, $79 a month, which appears nowhere in the query and had to come from the page. Neither sentence exists on our site in that wording. Google constructed both, from a source sentence with our name as the subject.

One of the six namings flatters us and should not. On the AI Mode question "what percentage of b2b contact data goes stale each year", the answer names Cleanlist inside a link label pointing at our data-decay post, alongside a link to ZoomInfo. That is a source credit rather than a shortlist slot. It counts in the 6 because a string search counts it, and a buyer reading that sentence is being sent to read a statistic rather than to consider a vendor.

What are the three different ways a page can be a ghost donor?

The 21 pages fail in three distinguishable ways, and the fix differs for each. This part transfers to any category.

A. Shape competition. The page carries plenty of compact, drop-in vendor lines and almost all belong to somebody else. Our enrichment listicle is the extreme case at 1 of ours to 18 rivals, cited 11 times and naming us twice. The B2B contact database page is 2 to 15, cited 3 times and naming us zero. These pages are not failing to write the shape. They write it for other people.

B. No vendor sentence exists at all. A definitional page answers the question and never says who does the job. /glossary/email-append was cited twice and named zero times, with no "Best for" construction of any kind on it, ours or anyone's. The answers took our definition and got their vendor list from three other providers. A page like that is a donor by construction.

C. No shortlist slot in the answer. The query is about one rival, the answer recites that rival's facts, and there is no slot for a second vendor. The RocketReach pricing guide, the Clearbit pricing guide and the Apollo versus ZoomInfo comparison all fail this way. The RocketReach guide has three correctly shaped Cleanlist sentences on it. They sit at character 20,985 and later, and the answer had no reason to reach for any of them, because it was not writing a shortlist.

Type C argues against a mechanical fix. You can add a perfect shaped sentence to a page whose answers are structurally single-vendor and get nothing back, because there was never a bullet for you to fill.

What did the independent recount change?

This study was written from a first-pass diagnostic that a second pass then re-derived from source. Six things changed, three of them materially, and all six are printed here because a study whose corrections live in a private file is asking to be trusted rather than checked.

1. The diagnostic read the wrong collection, which was the largest error. Its page list came from the superseded first pull of the census, the one that stored rank_absolute where it should have stored rank_group. The corrected pull is the collection of record. Reconciliation is conclusive: the superseded pull reproduces the diagnostic's per-page columns exactly and the corrected pull does not.

2. Naming counts roughly halve once the pull is fixed. Across the 39 page-citation pairs, 12 were published as named, 11 on a recount of the superseded pull, and 6 on the corrected one. Three rows move, all downward: the enrichment listicle from 36.4% to 18.2%, the Salesforce guide from 33.3% to 0%, the database cleaning post from 50% to 0%.

3. The universe was understated. The diagnostic reported 24 citing answers. The corrected figure is 37 answers across 21 pages and 39 page-citation pairs. The diagnostic covered 9 of those 21 pages and 27 of the 39 pairs.

4. The seven-times Cognism claim does not reproduce, for the segmentation reason above. Its companions, ZoomInfo 6, Clearbit 6 and Lusha 4, carry the same defect. Brand-level counts at Dice 0.30 are Cognism 12, ZoomInfo 11, Clearbit 8 and Lusha 2 answers.

5. The +0.68 correlation was withdrawn rather than corrected, for the reasons in the section above.

6. The source corpus was contaminated with code, and the corpus moved during the measurement. The first pass pulled "source sentences" out of our TypeScript data files that contained field names and metadata, so a matched sentence could read ..., primaryKeyword: "Clay alternatives", tldr: "Cleanlist is our pick.... The recount extracted string-literal contents only. Separately, six of the nine pages were edited in the working tree between 05:19 and 05:22, after the first pass wrote its output at 05:10, and one now opens with the diagnostic's own proposed sentence. Everything here was read through git show HEAD: so the corpus measured is the corpus that was live when the answers were generated.

The corrected per-page table for the nine pages with the most citations:

PageCitedNamedRateOur shaped linesRival shaped linesCleanlist mentions
/alternatives/clay33100%0031
/blog/2026-03-31-best-data-enrichment-tools-202611218.2%118154
/blog/2026-04-20-best-b2b-contact-database-software300%215126
/blog/2026-03-19-salesforce-data-enrichment-guide300%0025
/blog/2026-05-12-best-database-cleaning-software200%1831
/glossary/email-append200%0037
/blog/2026-03-07-apollo-vs-zoominfo100%005
/blog/2026-03-19-rocketreach-pricing-guide100%3291
/blog/2026-03-19-clearbit-pricing-guide100%1226

The top two rows are the whole argument and also its weakest point. /alternatives/clay is 3 citations, the enrichment listicle is 11, and everything else rests on 1, 2 or 3 observations each.

What are the limitations of this study?

Run-to-run instability is the largest single source of uncertainty here. The same 240-query census was collected twice on September 7, 25 minutes apart, identical query set, location, device and endpoint. Across the 27 nine-page citations common to both, one pull gives 9 namings and the other gives 5. Across the full census, naming stability for Cleanlist, defined as the share of answers naming us in either pull that name us in both, is 41.7%. Incumbents are stable (Cognism 94.7%, ZoomInfo 88.4%) and marginal brands flicker. At our level of corroboration a single-snapshot naming rate is close to a coin flip, so every naming rate on this page should be read as a range.

The sentence match is an inference and never a provenance signal. The random-page control fires on 24.8% of sentences at the 0.30 floor, so a meaningful share of the 269 would trace equally well to a page the answer never cited. The depletion result survives that objection because it compares against a base rate. The raw ratio does not.

n is 37 answers and 21 pages, one night, one country, one device, one category, with 11 citations on one page and 3 on another carrying the argument. The query set is ours, not a random sample of Google.

Sentence segmentation is a judgement call that already moved a headline once. The seven-times claim died on a splitter that would not break at a heading. Ours breaks at headings, list items, table cells and sentence terminators and drops anything under 38 characters. Change those choices and the integers change. The share does not.

The vendor attribution in the dataset is mechanical. A source sentence is attributed to the first brand name inside it, or failing that to the brand named in the nearest preceding heading. That is how a line reading only "Diamond Data: phone-verified mobile numbers with 98% connect rate" is attributed to Cognism, and it will be wrong on some rows.

Nothing here was validated by a fix. Five of the nine most-cited pages already carried a Cleanlist-subject "Best for" construction on the night of the measurement and three of those five were named zero times. This study identifies a pattern. It does not demonstrate that changing the pattern changes the outcome, and we will not claim it does until a re-measure says so.

Where can I download the dataset?

The open dataset is at /data/lifted-sentence-trace-2026-09.csv, published under CC BY 4.0. Attribute it to Cleanlist and link back to this page.

It is one file with two sections. Section one is 269 rows, one per traced answer sentence, sorted by source page then query. After a blank line, a marker row reading ## SECTION 2 introduces a 21-row per-page table with its own header.

Section one columns. query, the search string or AI Mode question. surface, either ai_overview or ai_mode. answer_sentence, the sentence as Google wrote it, citation markup stripped. matched_source_sentence, the highest-similarity sentence on the cited page, read at git HEAD. source_page, the cleanlist.ai path. similarity, the Dice coefficient, minimum 0.30. source_names_cleanlist, 1 or 0, and summing this column gives 16. brand_described, the brand the source sentence is about by the mechanical rule in the limitations, or none.

Section two columns. source_page. citing_answers. answers_naming_cleanlist, summing to 6. naming_rate_pct. traced_sentences, summing to 269. traced_from_cleanlist_naming_source, summing to 16. cleanlist_shaped_lines and rival_shaped_lines, summing to 9 and 45. cleanlist_mentions and rival_brand_mentions_all, word-boundary counts over the extracted page text. subject_brand, the brand the page is about, pipe separated where there are several. rival_brand_mentions_excl_subject and rival_mentions_per_cleanlist_mention_excl_subject, which is the density figure quoted above. shaped_sentence_note, a plain-language description of what shaped sentence the page did or did not carry, filled in for the nine most-cited pages.

To reproduce it without us: pull any Google SERP API for the query list with location_code 2840, English, desktop, requesting the AI Overview element with the asynchronous flag set. Keep the answers whose body or reference list contains your own domain. Fetch each cited page at a commit rather than from the live site. Segment both sides at headings, list items, table cells and sentence terminators. Score each answer sentence against each source sentence with Dice over four-character word tokens and keep the best match at 0.30. Then count how many of the winners contain your own brand name. Expect your absolute counts to differ from ours, because these answers are regenerated per request and the naming layer is unstable. What we would expect to survive is the shape of the result.

What should a marketer do on Monday?

This applies to anyone publishing a comparison listicle, a category roundup or an alternatives page, in any category. The mechanism has nothing to do with B2B data. It has to do with writing eighteen ready-made answer bullets and putting your name on one of them.

1. Count the drop-in lines on your page, and count whose name is the subject. Search for "Best for" and for every construction of the form Brand: <short verdict>. For each hit, ask whether the brand is the grammatical subject of a sentence that would survive being copied out of the page on its own. That is a drop-in answer bullet. Our ratio across 21 pages was 1 of ours to 5 of theirs, and 1 to 18 on our most-cited page. Half an hour with a text search will tell you yours.

2. Write one self-contained sentence with your brand as the subject. Not "the best tool for X is Brand", and not "alternatives like Brand". Subject position, a narrow buyer segment, a mechanism containing a number, and a real price with a period. Our best-performing page does not use the literal "Best for" string at all. It leads with a summary sentence whose subject is us, whose object is a specific buyer, and which contains a price. Google built the bullet out of that by itself.

3. Do not cut the competitors. This is the counter-intuitive result and it is worth holding on to. Rival mention density does not predict naming in either direction here, so deleting rival coverage costs you the reason the page ranks and buys nothing measurable. Change the shape of the coverage instead: keep every fact, and write the rival's verdict as a clause inside a longer sentence rather than as a standalone bullet that copies cleanly.

4. Never give yourself a $0 or "free" headline price in a summary. Price is the field these answers quote most reliably. Our Clearbit pricing guide gave our own price as "$0/mo" in its summary and was named in 0 of 1. The page that got named carried $79 a month, and the answer repeated it.

5. Check that your roster section is not empty. One of our two near-identical listicles has a summary section under the heading listing the tools reviewed, and the other has an empty section that jumps straight to the first entry. The first was named twice, the second zero times. That is one observation and not a law, and filling an empty summary section costs ten minutes.

6. Do not chase "put it in the first 1,500 characters" on this evidence. The one bullet Google demonstrably re-emitted, with our price attached, sits nearly 29,000 characters into the page.

7. Measure twice before believing a naming number, yours or a vendor's. Two identical pulls 25 minutes apart moved our naming count from 9 to 5 on the same 27 citations. Any single-snapshot claim about which brands an AI answer names, including several on this page, is one coin flip wide.

8. Trace rather than guess. The method is four steps and no machine learning: take the answer text, split it into sentences, match each one back to the page the answer cited, then count how many of the winners contain your own name. If that count is small, you are writing somebody else's answer. Ours was 16 out of 269.

What is Cleanlist doing about its own result?

Publishing it, then fixing the pages, in that order. The work is a shaped, self-contained sentence with Cleanlist as the subject on each cited page that lacks one, and a rewrite of rival entries so their coverage survives intact while their sentences stop being drop-in answer bullets. No competitor is dropped, no fact is deleted, and no partner is repositioned as a competitor.

Then we re-measure on the same query set, at least twice in the same session, because a single pull cannot tell the difference between a fix and a coin. We will publish the result including if the number does not move.

References & Sources

  1. [1]
  2. [2]
  3. [3]

Enrich your own contacts free for 14 days

Get verified emails and direct dials back on your own contacts and export them to your CRM or sending tool. 1 credit per verified email, 10 per direct dial. 250 credits, 3 seats, 14 days. No card required.

Start free trial

250 credits, 3 seats, 14 days. No card required.

Run it on the list you already have.

Cleanlist puts one lookup through 25+ providers and stops at the first source that returns. Search is free and unlimited on every plan. A verified work email is 1 credit, a direct dial is 10, both together are 11, and a miss costs nothing at all.

14-day Scale trial: 250 credits, 3 seats, no card, every feature except the public API and MCP. The Free plan stays at 30 credits a month after that.