There are three Reddit threads about waterfall enrichment right now: one asking if it’s worth it, one in r/snowflake asking if it’s just Clay influencer bait, and one that explains what it actually does behind the scenes. All three are right about different parts of it, which is exactly why the topic is confusing.

Here is the honest version, from building waterfall enrichment for client programs at KomsGro: it works, it saves you money per verified contact, and most implementations waste 40 percent of their credits on providers that were never going to find the data anyway.

What waterfall enrichment actually does

The concept: instead of running every contact through one enrichment provider (and paying their rate whether they find the data or not), you sequence providers cheapest-first and only escalate to the next when the current one fails.

The mechanics are simple:

  1. Provider 1 (cheapest, broadest coverage): query the contact. If it returns an email or phone, done.
  2. Provider 2 (mid-price, deeper): query the contact again. If it returns data, done.
  3. Provider 3 (most expensive, highest accuracy): last resort for the contacts the first two missed.

The reason this saves money: Provider 1 might find 60 percent of contacts at $0.05 each. Provider 2 catches 30 percent of the remainder at $0.20. Provider 3 catches 10 percent of what’s left at $1.00. The blended cost per verified contact drops from a flat $0.50 (one premium provider) to something like $0.12.

That is the theory. In practice, the waterfall leaks value in three ways that most teams do not account for.

The three leaks

1. Duplicate queries across providers. If you query Provider 1 and Provider 2 with the same contact (instead of only escalating failures), you pay twice for the same data. The waterfall needs a gate: only escalate contacts that returned nothing.

2. Verifying outputs from cheap providers. A cheap provider returns an email, but is it deliverable? If you skip verification and it bounces, you have paid for a bad row and poisoned your sender reputation. Always verify the waterfall output before it touches a mailbox.

3. Rate limits and credit caps. If you run the waterfall across 1,000 contacts and Provider 2 is slow, the workflow may time out and re-queue rows that already succeeded. Idempotency (checking the enrichment log before escalating) prevents duplicate spend.

The provider order (practical, not theoretical)

The providers I use, cheapest to most expensive, for a B2B contact waterfall:

  1. The CRM’s own enrichment. If your CRM (HubSpot, Attio) has built-in enrichment, use it first. It is usually free or bundled.
  2. A broad data provider (Apollo at its cheapest tier). Covers the widest range but with the most stale data.
  3. A specialized email finder (Prospeo, Findymail, Anymail Finder). Higher accuracy, higher price per contact.
  4. A mobile finder (if you need phone numbers). The most expensive layer, and the one with the lowest hit rate.

The waterfall does not need to be four providers deep on every campaign. Two layers (CRM + one specialized finder) cover most contacts at the lowest blended cost. Add layers only when the coverage rate drops below what your volume needs.

The n8n implementation

The workflow architecture in n8n:

Trigger (new contacts batch)
  -> CRM enrichment check (skip if already enriched)
  -> Provider 1 API call
  -> IF email found -> verify -> write to CRM -> done
  -> IF no email -> Provider 2 API call
  -> IF email found -> verify -> write to CRM -> done
  -> IF no email -> Provider 3 API call (optional)
  -> IF still nothing -> flag as "unreachable", skip
  -> Log: rows in, rows enriched by each layer, rows unreachable

Two node-level rules: the enrichment log write happens before the next provider call (not after), so a timeout does not cause a re-query. And the verification step runs on every waterfall output, regardless of which provider produced it.

The blended cost math

Here is a real-shaped example for a 500-contact list:

  • Layer 1 (CRM enrichment, free): 40 percent coverage = 200 contacts at $0
  • Layer 2 (Apollo, $0.03/contact): 150 more contacts = $4.50
  • Layer 3 (Prospeo, $0.20/contact): 80 more contacts = $16.00
  • Layer 4 (mobile finder, $1.00/contact): 30 contacts = $30.00

Blended cost: $50.50 for 460 verified contacts = $0.11 per verified contact, versus $250 if every contact went through the mobile finder. That is the waterfall math, and it is why the approach is real even if the marketing around it is overhyped.

When waterfall enrichment is not worth it

Three cases where I tell teams to skip the waterfall:

  1. Your list is under 100 contacts. The setup cost (building the workflow, configuring each provider) exceeds the savings on a small batch.
  2. You only need emails, not phones. Two providers cover email at a blended rate that a single broad provider matches at volume.
  3. Your data providers have overlapping databases. If Provider 2 uses the same source as Provider 1, escalating is a waste. The waterfall only works when each provider has genuinely different coverage.

What to do this week

If you are already paying one enrichment provider flat-rate, the waterfall is the single highest-ROI improvement you can make to the stack. Build the two-layer version (CRM first, one specialized finder second) in n8n, run it on your next list, and compare the blended cost per verified contact against your current flat rate.

If you want the whole system built (waterfall, dedupe, verification, CRM sync, and sending), that is KomsGro’s outbound marketing service. The waterfall sits inside the GTM engineering stack between the data layer and the sending layer.

The coverage problem (and why the waterfall exists)

The reason waterfall enrichment exists is simple: no single data provider covers every contact. Apollo has the biggest database but the most stale entries. Prospeo is more accurate but covers fewer companies. Lusha is strong on phone numbers but weak on emails. No single provider returns usable data for more than 60 to 70 percent of contacts.

The waterfall approach fixes this by running contacts through providers in sequence until one returns usable data. The coverage gain is real: two to three providers typically reach 75 to 85 percent total coverage, versus 50 to 65 percent from any single provider. The cost logic is also real: if the cheap provider covers most contacts, the expensive layers only process the remainder.

The deduplication layer (the part that saves the most money)

The most expensive bug in a waterfall is a contact that runs through all providers and returns data from all of them. You pay three times for the same email, and you have to decide which one is right.

The fix is a deduplication check between every layer: before escalating to the next provider, check whether the previous layer already returned a verified email. If it did, skip the escalation. This is a single IF condition in n8n, and it is the highest-ROI line of logic in the entire workflow.

The verification layer (never skip it)

Every waterfall output must pass a verification step before it enters a campaign. The reason is that a waterfall catches contacts from multiple providers, and some of them will be stale, disposable, or role-based addresses that bounce.

Verification tools (ZeroBounce, NeverBounce, or the verification built into your sender) cost $0.001 to $0.005 per contact and catch most bounces before they happen. The trade is trivial: a penny of verification against a burned domain and weeks of reputation repair.

The tech-stack waterfall (a variant worth knowing)

The same waterfall principle applies beyond contact data. You can build a technographic waterfall: query multiple tech-stack databases (BuiltWith, TheirStack, Wappalyzer) in sequence until one confirms the technology you are targeting. The logic is identical to contact enrichment: cheapest-first, escalate on failure, dedupe.

This variant is useful when you are building a list of companies that use specific technologies (which is covered in how to find leads by technology stack). The waterfall principle is the same: maximize coverage, minimize cost per verified result.

The waterfall in practice (what a run looks like)

To make this concrete, here is what a waterfall run looks like for a single contact in n8n, with the actual nodes:

  • Node 1: dedup check. Look up the contact email in the enrichment log. If found with a verified email, skip everything and return the cached result. This one condition is the highest-ROI line in the workflow.
  • Node 2: CRM enrichment. Call the CRM’s own enrichment endpoint. If it returns an email, verify it and write the result. If not, escalate.
  • Node 3: Apollo API call. Standard tier, broadest coverage. If it returns an email, verify it.
  • Node 4: IF check. Did Apollo return a valid email? If yes, verify and write. If no, escalate.
  • Node 5: Prospeo API call. Specialized email finder. Higher accuracy, higher price.
  • Node 6: IF check. Found? Verify and write. If not, flag as unreachable and skip.
  • Node 7: verification call. Run the email through a verifier (ZeroBounce or the sender’s built-in check). A verified email writes to the CRM; a bounced or disposable address gets flagged and excluded.
  • Node 8: log write. Every run writes: contact email, layers used, cost incurred, result (enriched or unreachable).

The entire workflow is eight n8n nodes, and it replaces the flat-rate enrichment subscription with a pay-per-verified-result model that saves 40 to 60 percent on data costs.