How Waterfall Enrichment Works

CleanlistThe short answer

Waterfall enrichment is a method that runs one data lookup across several B2B data providers in sequence, stopping at the first provider that returns an answer the system can verify. Rather than trusting one database to hold every contact, the record falls from provider to provider until one of them answers, which is why a waterfall returns more coverage than any single source inside it. Cleanlist walks a pool of 25+ providers on every lookup, verifies each candidate work email at the mailbox before accepting it, and charges once for the answer, 1 credit for a verified work email and 10 for a direct dial, however many providers had to run to find it.

  1. 01What is waterfall enrichment?
  2. 02How does a waterfall enrichment run actually execute?
  3. 03Why does waterfall enrichment beat a single data provider?
  4. 04What does waterfall enrichment cost per record?
  5. 05Which data providers are in the Cleanlist waterfall?
  6. 06How is waterfall enrichment different from licensing a contact database?
  7. 07Does the waterfall verify data, or only find it?
  8. 08What happens when the waterfall finds nothing?
  9. 09Does waterfall enrichment work for company data as well as contacts?
  10. 10Can you build a waterfall enrichment sequence yourself?
  11. 11How large a list can you run, and does it run synchronously?
  12. 12Where does waterfall-enriched data land?
  13. 13How do you evaluate a waterfall enrichment vendor?
  14. 14When is waterfall enrichment the wrong tool?

What is waterfall enrichment?

Waterfall enrichment is one lookup walked across several data providers in sequence, stopping at the first provider that returns a verified answer. Cleanlist walks a pool of 25+ providers per lookup and charges once for the answer, however many providers had to run to find it.

The name describes the shape of the run. A record enters at the top, the first provider is asked for it, and if that provider holds nothing the record falls to the second, then the third, and so on until either a provider returns a value that passes verification or the pool is exhausted. Nothing about the sequence is visible to the person who asked for the contact. They upload a list and get a list back.

The reason the pattern exists is coverage. No B2B data provider holds every person in every market, and the gaps are not random: a provider that is strong on North American technology companies is often thin on European manufacturing, and a provider that is strong on senior titles is often thin below director level. Asking one provider is asking one opinion. Asking several in order, and stopping at the first that answers, is what turns several partial datasets into one usable one.

Waterfall enrichment is a routing strategy rather than a dataset. It describes how a lookup is executed, not who owns the underlying record, which is why the same term covers a team wiring four provider APIs together in a script and a platform that ships the sequence already assembled.

The same idea travels under several names, and they mean the same thing in practice: cascading enrichment, multi-provider enrichment, multi-source enrichment, provider fallback and, inside engineering teams, simply a fallback chain. If a vendor says it queries providers in priority order and bills you for the result rather than the attempt, it is describing a waterfall whatever it calls it.

How does a waterfall enrichment run actually execute?

Each field runs its own waterfall. Cleanlist normalises the input, orders the pool in cost order, calls providers one at a time, verifies each candidate answer before accepting it, stops at the first that passes, and bills once for the field.

Email and phone are separate runs on the same person, because the providers that are good at one are frequently not the providers that are good at the other. A record can therefore come back with a verified work email found on the second provider asked and a direct dial found on the ninth, or with an email and no phone at all.

Cost order matters more than it sounds. Running the cheapest provider that could plausibly hold the record first, and only escalating when it comes back empty, is what keeps a per-result price viable. A sequence that always calls the most expensive provider first would find the same answers and cost several times as much to operate.

Verification sits inside the loop rather than after it. A provider returning a string that looks like an email is not a hit. The address is checked against syntax, then the domain's mail exchange records, then the mailbox itself, and only an address that survives that is accepted and billed.

The four moving parts, named:

Input normalisation. Names, domains and company strings are cleaned to a common shape before any provider is called, so the same person is not missed by one provider over a formatting difference.

Provider order. Cost order, set by Cleanlist rather than by the buyer, so a lookup escalates only when the cheaper source came back empty.

Stop condition. The first verified answer. Providers below the hit are never called, which is why the price is the same whether the answer came from the first provider or the last.

Billing. Once per accepted field. 1 credit for a verified work email, 10 for a direct dial, 11 for both on the same contact.

Why does waterfall enrichment beat a single data provider?

Because the gaps in one provider are not the gaps in the next. On the Cleanlist 500-Lead Enrichment Benchmark, 2026, a 25+ provider waterfall returned 98% verified work email and 85% direct dial across 500 stratified B2B leads, against 70% to 80% email and 30% to 60% phone from single sources on the identical input.

The interesting half of that result is the phone column. Email coverage from a good single provider is already decent, so the waterfall buys perhaps twenty points there. Direct dials are the field where single sources fall off a cliff, and the field where a sequence of providers changes what a calling team can actually do with a list.

The comparison was run the honest way: one list, run once through the full waterfall and once through single sources, with the same 500 leads going into both. Coverage claims that do not name their denominator or their input list are not comparable to each other and should not be treated as if they were.

There is a second effect that a coverage number does not show. A record that only one provider in the pool holds is a record your competitors calling the same market probably do not have, because they are asking one provider and it is not that one.

What does waterfall enrichment cost per record?

Cleanlist charges per accepted field rather than per provider call: 1 credit for a verified work email, 10 for a direct dial, 11 for both, 0.5 for validating an address you already hold, and 0 when the whole pool comes back empty. People Search and Company Search cost nothing at all.

Because the price is per result, the arithmetic is the same every month. Starter is $79 for 1,500 credits, which is 1,500 verified work emails or about 136 fully enriched contacts at 11 credits each. Pro is $229 for 5,000 credits and Scale is $599 for 15,000. Annual billing takes 25% off every paid plan, and an extra seat is $20 a month.

The line worth reading twice is the last one. A lookup that walks all 25+ providers and finds nothing costs zero credits. That is the difference between paying for data and paying for attempts, and it is why a pool this deep is affordable to run: escalation is free to the buyer and expensive only to us.

Credits come out of one shared team wallet on every plan, including Free, so a rep who runs four hundred lookups and a rep who runs four spend from the same balance and nobody buys a second subscription to try it.

Which data providers are in the Cleanlist waterfall?

More than 25, and the pool includes Wiza, Prospeo, Findymail, LeadMagic, Datagma, Crustdata and Lusha alongside the rest. Cleanlist holds the contracts, the API keys and the rate limits, so a customer buys one plan rather than 25.

The commercial argument for a managed pool is not that any one of those providers is bad. It is that a team wiring them together itself signs 25 agreements, holds 25 sets of credentials, reconciles 25 invoices against one month of usage, and rebuilds a parser every time one of them changes a response shape. That work never finishes, and it is not the work a revenue team was hired to do.

The sequence is ours to set and to keep tuned. Provider coverage moves, prices move, and a provider that was the right first call for European mobiles last quarter may not be this one. Tuning the order centrally is the part of the product that keeps working after the integration is done.

How is waterfall enrichment different from licensing a contact database?

A licensed database sells you access to one company's stored records for a fixed annual fee. A waterfall spends nothing until a lookup is asked for, queries several sources at that moment, and verifies the answer before it is delivered.

The two models fail differently. A database's weakness is freshness: a record sits in it between refresh cycles, and B2B contact data decays at roughly 2.1% a month and 22.5% a year according to Cognism's published figures, with email addresses the fastest field to go. A waterfall's weakness is that it answers only the question you asked. You cannot browse it the way you can browse a database you licensed.

The cost shapes are different too. A licensed database charges the same whether a seat used it heavily or not at all, which is why per-seat data spend tends to drift away from usage over a year. A waterfall charges per returned record, so the invoice tracks what the team actually looked up.

Neither model is universally right. A team that wants to sit inside a UI and explore a market of 300 million stored contacts is buying a database, and should. A team that has a list, a CRM or a search result and wants the missing fields filled and verified is buying a waterfall.

Does the waterfall verify data, or only find it?

It verifies before it accepts. Every candidate email is checked against RFC 5322 syntax, then the domain's MX records, then the mailbox itself over an SMTP handshake, with nothing sent to the address, and an address that fails is not counted as a hit and not billed.

Risk is reported rather than hidden. A catch-all domain that accepts every address is returned flagged as accept_all instead of green, a role address such as info@ or sales@ is labelled as one, and a disposable domain is marked and dropped. Collapsing those three into a single green tick is the industry habit this product exists to argue against, because each of them needs a different decision from the person about to send.

Validating a list you already hold is the same machinery at 0.5 credits an address, without the enrichment. That is the cheaper first move for a team whose problem is bounce rate rather than coverage.

The verdicts, and what each one means for the person about to send:

Deliverable. The mail server confirmed the mailbox exists. Safe to send.

Catch-all. The domain accepts everything, so a yes proves nothing. Flagged, never guessed.

Invalid. The mail server refused the address. Never charged, never synced.

Disposable. A burner domain. Marked and dropped.

Role address. info@, sales@ and the rest. Labelled so you can route it rather than treat it as a person.

What happens when the waterfall finds nothing?

The row comes back marked as a miss and costs zero credits. No provider in the pool is billed to the customer for an empty answer, and no placeholder or pattern-guessed address is written in to make the run look complete.

That last part is the one to check when comparing tools. A guessed address in the shape firstname.lastname@company.com is trivial to generate and will pass a syntax check, so a vendor that fills gaps with guesses can report a coverage number close to 100%. It shows up later as bounces, and bounces are charged to your sending domain rather than to the vendor.

A miss is also information. On a list where a quarter of the rows come back empty, the useful next question is usually about the list rather than about the enrichment: the ICP may be aimed at companies too small to have public contact data, or at a market where the pool is genuinely thin.

Does waterfall enrichment work for company data as well as contacts?

Yes, and the firmographic side behaves differently enough to be worth separating. Company fields are attributes of an organisation rather than of a person, they decay far more slowly than an email address, and on Cleanlist an enriched record comes back with up to 180 firmographic properties while Company Search itself costs 0 credits.

The waterfall still matters here, because company coverage is uneven in the same way contact coverage is. A provider with good headcount and industry data on North American software companies is often weaker on privately held European manufacturers, so the same sequence-and-stop pattern applies to a field like employee count.

What separates the two in practice is the refresh cadence. A person changes jobs and their work email dies the same day, which is why contact fields need reverification. A company's founded year never changes, its industry rarely does, and its headcount moves on a quarterly rather than a daily scale. A team that reverifies contacts monthly and firmographics quarterly is usually spending in the right proportion.

The boundary to know before you buy: the stored company record is the descriptive set, name, domain, headcount, industry, headquarters location and founded year among the rest. Signals about a company's current situation, funding events, hiring activity, tech stack or recent news, are produced by an AI research column at the moment you ask for them rather than read out of a stored field, and they are priced separately.

Can you build a waterfall enrichment sequence yourself?

Yes, and teams with an engineer to spare do. The build is a weekend and the maintenance is permanent, which is the part that decides it.

The moving parts are the same for everybody: a commercial agreement and a key per provider, a normalisation layer so a name and a domain reach every API in the shape it expects, per-provider rate limiting and retry, a parser per response format that has to be repaired whenever a provider ships a change, a verification stage so a returned string is not trusted on sight, deduplication when two providers answer with different values for the same field, and a reconciliation job that matches 25 invoices against one month of internal usage.

The honest test is whether somebody owns it as part of their job. A sequence nobody owns drifts: a provider quietly goes empty for a region, the coverage number falls, and the first sign is a rep complaining that the data got worse.

There is a real case for building. If enrichment is inside your own product rather than inside your go-to-market, if you have negotiated volume rates a reseller cannot match, or if your legal position requires direct contracts with each data source, the build is the right answer and the maintenance is a cost you were always going to carry.

How large a list can you run, and does it run synchronously?

A list holds 2,500 leads on any paid plan and 100 on the Free plan, and a bulk run is asynchronous. Enrichment jobs return a workflow_id that you poll until the results are ready, rather than holding a request open.

Rows enter a list from a CSV upload, a CRM import, People Search and Company Search, the Chrome extension, or LinkedIn and Sales Navigator, so a list does not have to start life in a spreadsheet. CSV upload is a paid-plan feature, and CRM import arrives with the two-way sync on Pro and Scale.

Asynchronous is the right shape for a waterfall specifically. A single lookup that escalates through several providers takes as long as the providers it had to ask, so a synchronous API would spend most of a large run holding connections open and timing out on exactly the records that needed the most work.

Where does waterfall-enriched data land?

Into a Cleanlist lead list first, then out. Two-way sync with HubSpot, Salesforce and Pipedrive on Pro and Scale, one-way out to Outreach, Salesloft and Lemlist, CSV export on every plan, plus the REST API and the MCP server. Cleanlist does not send email.

Pushing a finished lead to a CRM costs 0.2 credits. The public REST API is a Pro and Scale feature and authenticates with clapi_ Bearer keys and OAuth scopes, and every request returns a signed cost quote before any credits are spent, so an integration can decline a job it thinks is too expensive before it runs. The MCP server is available from Starter upward, which is what lets an assistant such as Claude run enrichment directly. Both are held back during the 14-day trial.

The direction of each integration is worth stating plainly, because it is the thing buyers assume rather than check. HubSpot, Salesforce and Pipedrive read and write. The three sequencers receive rows and send nothing back. Sending the email itself is somebody else's product.

How do you evaluate a waterfall enrichment vendor?

Run your own list through it and count, because every published coverage number was measured on somebody else's input. Take 200 to 500 real records from your own ICP, run them through each vendor unchanged, and compare four things: how many rows came back with a work email, how many with a direct dial, what each run cost, and how many of the returned emails survive an independent verification pass.

That last check is the one people skip and the one that separates vendors. Export the returned addresses, verify them somewhere other than the tool that produced them, and see how many hold up. A vendor that pattern-guesses gaps will post a high fill rate and a low survival rate, and the difference is bounces you pay for later.

Five questions worth asking before the test, because the answers change how you read the result:

Are misses billed? If a lookup that returns nothing still costs money, a deep pool becomes expensive rather than free to escalate, and the vendor has an incentive to guess.

Is billing per provider call or per accepted field? Per call means the price of a record varies with how hard it was to find, which makes a budget impossible to forecast.

Is verification inside the run or sold separately? An address that was never checked at the mailbox is a candidate, not a result.

Are catch-all and role addresses reported honestly? A vendor that returns a catch-all as green has moved a decision onto you without telling you.

Who tunes the provider order, and how often? A pool that was tuned at signup and never again decays quietly as provider coverage shifts.

One caveat about our own position. Cleanlist does not hold SOC 2 Type II or ISO 27001 certification, and if your security review requires either of them, that is a real reason to buy elsewhere and worth finding out on day one rather than in procurement.

When is waterfall enrichment the wrong tool?

When the field you need is not a contact field. A waterfall finds and verifies who somebody is and how to reach them. It does not tell you that they are in a buying cycle, what software they run, or what happened on their last earnings call.

Three cases where a different purchase is the right one. If your motion depends on intent signals, a platform built around them, such as ZoomInfo or Cognism, is what you are actually shopping for. If you want to compose your own multi-step research pipelines and have somebody who will own them as a job, Clay is built for that. If you need to browse and slice a large stored database as a research surface rather than fill gaps in a list you already have, license a database.

There is also a boundary inside Cleanlist worth knowing before you buy. The stored fields are the contact and company record: work email, direct dial, LinkedIn URL, title, seniority, department, company name, domain, headcount, industry, headquarters location and founded year. Anything past that set, funding events, hiring signals, tech stack or recent news, is produced by an AI research column at the moment you ask for it rather than read out of a stored field, and it is priced separately.

The shortest way to describe the boundary: Cleanlist is an orchestration layer over other people's data rather than a data vendor holding a database of its own.

The follow-up questions.

Who is waterfall enrichment for?

Teams that already have a list and need the missing contact fields filled and verified: outbound sales teams building call and email lists, RevOps teams keeping a CRM current, and growth teams enriching inbound signups. It is a poor fit for a team that has no list yet and wants a research surface to browse, and a poor fit for a team whose bottleneck is intent rather than contactability.

Do I need a credit card to try waterfall enrichment?

Not at Cleanlist. The live offer is a 14-day Scale trial with 250 credits and 3 seats and no card required, which is enough to enrich roughly 22 contacts with both an email and a direct dial, or 250 with an email alone. The Free plan includes 30 credits a month afterwards, with a 100-lead list limit and no CSV upload.

Does waterfall enrichment work outside North America?

It works anywhere the pool has coverage, and coverage varies by region rather than being uniform. A pool of 25+ providers exists precisely because regional strength is uneven: providers differ sharply on European mobiles, on APAC contact data, and on non-English company records. The practical test is cheap, because a lookup that finds nothing costs zero credits, so running a few hundred records from the region you actually sell into tells you more than any published coverage figure.

What is the difference between waterfall enrichment and data appending?

Data appending is the outcome, waterfall enrichment is one way of producing it. Appending means adding missing fields to a record you already hold, whatever the source. A waterfall describes how those fields were sourced: several providers queried in order, with the first verified answer accepted. An append done against a single licensed file and an append done through a waterfall produce the same column in your CRM and very different fill rates.

Does waterfall enrichment reduce bounce rates?

It reduces bounces caused by wrong or dead addresses, and it does nothing about bounces caused by your sending setup. Enrichment that verifies at the mailbox will not hand you an address the mail server refuses, and catch-all domains are flagged rather than passed off as deliverable. Bounces that come from a missing SPF or DKIM record, a cold domain, or volume ramped too fast are a deliverability problem, and no data vendor can fix them for you.

Can an AI assistant or an agent run waterfall enrichment for me?

Yes, through the MCP server, which is available from Starter upward and lets an assistant such as Claude search for people and companies and enrich a list directly. The REST API is the other route and is a Pro and Scale feature, authenticating with clapi_ Bearer keys and returning a signed cost quote before any credits are spent. Both are held back during the 14-day trial.

Does Cleanlist hold SOC 2 or ISO 27001 certification?

No. Cleanlist does not currently hold SOC 2 Type II or ISO 27001, and vendors such as Cognism and ZoomInfo do. If your security review makes either certification a hard requirement, that is a genuine reason to buy elsewhere. The providers in the waterfall are disclosed as subprocessors in the privacy policy, which is the document a data protection review will ask for first.

Can I choose the order the providers run in?

No, and that is deliberate. Cleanlist sets and maintains the cost order centrally, because provider coverage and pricing move constantly and an order tuned once at setup decays. What you control is the field: email only at 1 credit, phone only at 10, or both at 11, so you decide what to spend on rather than which supplier answers first.

Gain full access for 14 days.

Cleanlist runs one lookup across 25+ providers and stops at the first source that returns. Search costs nothing on every plan, a verified work email is 1 credit, a direct dial is 10, and a miss costs nothing at all.

250 credits, 3 seats, 14 days. No card required. Every feature except the public API and MCP. The Free plan stays at 30 credits a month after that.