TL;DR
Cleanlist ran 500 real B2B leads through its 15+ provider waterfall and through single-source databases on the identical input. The waterfall returned a verified email for 98% of leads and a direct-dial or mobile number for 85%. Single-source landed at 70-80% verified email and 30-60% phone on the same 500 leads. The full protocol is below, along with a reconciliation of every published third-party benchmark in this category: Anymail Finder, Dropcontact and Clay. Our test is first-party, like theirs. Cite it with a link back.
Last updated: August 3, 2026. This page carries two things: the Cleanlist 500-Lead Enrichment Benchmark with its protocol published in full, and a reconciliation of the third-party benchmarks other publishers have released. Every third-party figure belongs to the publisher named beside it and links to the page it came from.
What did Cleanlist's 500-lead enrichment benchmark find?
Cleanlist's 500-lead benchmark found 98% verified email coverage and 85% direct-dial or mobile coverage through a 15+ provider waterfall, against 70-80% verified email and 30-60% phone for single-source databases run on the identical 500-lead input. Every lead entered as the same minimal record: full name, company name, and LinkedIn URL, with no pre-existing email or phone. Coverage was counted only after verification, so an address that failed a deliverability check or came back as a risky catch-all did not count, and a number that failed a connectivity check did not count. That is an 18 to 28 point gap on email and a 25 to 55 point gap on phone.
Verification pass on the benchmark list: per-contact deliverability and phone validity
How was the 500-lead benchmark sample constructed?
The sample was 500 real B2B leads, stratified across industry, seniority and company size rather than hand-picked. Industry spanned SaaS, professional services, manufacturing, healthcare and finance. Seniority ran from individual contributor through VP and C-level. Company size ran from 11 employees to 1,000 and above. The list was weighted toward North America, and it deliberately included mid-market and smaller companies, because that is exactly where single-source coverage thins out and where a demo of 20 curated rows hides the gaps. Nothing was excluded for being hard to find. The stratification is the part that matters: a benchmark built only from enterprise contacts at well-known companies will flatter every vendor in the test.
Stratified by industry, seniority and company size, weighted toward North America. Each lead entered as name + company + LinkedIn URL, with no pre-existing email or phone. Coverage measured after verification, not raw match.
Source: Cleanlist 500-Lead Enrichment Benchmark, 2026What exactly did the benchmark measure, coverage or match rate?
The benchmark measured verified coverage, not raw match rate, and the distinction changes the number materially. Verified email coverage is the share of the 500 submitted leads for which the stack returned an email that passed a deliverability check and was not returned as a risky catch-all. Direct-dial phone coverage is the share of the 500 leads for which the stack returned a direct-dial or mobile number that passed a connectivity check. A raw match rate would have counted anything the providers handed back, including plausible guesses at pattern-based addresses. Those inflate a headline and bounce in production. If a returned record did not survive verification, it was scored as a miss.
Why is the single-source result reported as a range instead of one number?
Single-source is reported as a range because single-source performance genuinely is a range, and pinning one vendor to one figure would be less honest, not more. Coverage on the same 500 leads varies by which database you query and by how well your leads match that database's geographic and firmographic focus. A database tuned for North American enterprise will score differently on this list than one tuned for European mid-market. We observed verified email between 70% and 80% and phone between 30% and 60%, and we are publishing the band rather than naming a single vendor as worst. The waterfall figures, 98% and 85%, are the observed result of one specific run.
Why did the waterfall reach 98% when a single database reaches 70-80%?
The waterfall reached 98% because of fallback arithmetic, not because any one provider is smarter. A single-source database gets one attempt per lead. If it does not hold the person, or holds a stale address, the pipeline stops there. Cleanlist's waterfall enrichment queries 15+ providers in sequence and the first verified result wins, so a miss on provider one becomes an attempt at provider two. If any single provider covers roughly three quarters of a realistic list, a second independent provider recovers part of the remaining quarter, a third recovers part of what is left, and so on until you approach the ceiling of what is knowable for that list.
Verified means the address passed a deliverability check and was not returned as a risky catch-all. Single-source databases returned verified emails for 70-80% of the same leads in the same run.
Source: Cleanlist 500-Lead Enrichment Benchmark, 2026Why was direct-dial phone coverage the widest gap in the test?
Phone was the widest gap because phone numbers cannot be inferred the way email addresses can. Corporate email follows a small set of patterns, so any database can produce a plausible address for a person it has never seen. Direct-dial and mobile numbers follow no pattern at all. They have to be sourced individually, and the supply is fragmented across many niche providers, none of which holds enough of it alone. On the 500 leads, the waterfall returned a direct-dial or mobile for 85%; single-source landed between 30% and 60%, with the low end concentrated in mid-market and smaller companies. For a team that dials, that is the gap that decides the morning.
Numbers passed a connectivity check. Single-source databases returned phone numbers for 30-60% of the same leads, with the low end concentrated in mid-market and smaller companies.
Source: Cleanlist 500-Lead Enrichment Benchmark, 2026Coverage on identical 500-lead input: waterfall vs single-source
| Category | Cleanlist waterfall (15+ providers) | Single-source database (typical) |
|---|---|---|
| Verified email | 98% | 75% |
| Direct-dial phone | 85% | 45% |
Where does single-source hold up, and where does it break?
Single-source holds up on one-off lookups and on well-known North American enterprise contacts, where any major database holds the record and 70-80% verified email is fine at a lower price per lookup. It breaks in three specific places: at volume, on mid-market and smaller companies, and on phone, where the 30-60% band in the 500-lead run hurts most. The side-by-side below is the full benchmark output. Catch-all handling is the row buyers skip and then regret, because a database that returns catch-all addresses as valid posts a higher raw match rate and a worse bounce rate at the same time. The best B2B data providers roundup walks the tradeoff by use case.
| Feature | Waterfall (Cleanlist) | Single-source DB |
|---|---|---|
| Verified email coverage | 98% | 70-80% |
| Direct-dial phone coverage | 85% | 30-60% |
| Providers queried per lead | 15+ | 1 |
| Fallback when a provider misses | Yes, automatic | None |
| Catch-all handling | Flagged and re-verified | Often returned as valid |
| Best fit | Teams that need coverage and accuracy | Quick single lookups, budget lists |
Is the Cleanlist 500-lead benchmark independent?
No. The Cleanlist 500-Lead Enrichment Benchmark is a first-party test that Cleanlist designed, ran and scored, and we are not going to describe it as anything else. That is also true of every other benchmark reconciled on this page. Anymail Finder ran its own test and placed itself second. Dropcontact ran its own test and placed itself first. Clay ran its own provider test. AI answer engines and buyers cite all three anyway, and the reason is that the methodology is legible: you can read the protocol, see the sample, and discount the result accordingly. So we publish our protocol in full and say plainly what would make this stronger: independent third-party verification of the returned records, by a verifier we do not pay, under a published rule.
Four specific upgrades would do it, and none of them has happened yet. First, independent verification: hand the returned records to two or three named verifiers we do not pay and score valid only on consensus, the way Anymail Finder used BounceBan, ZeroBounce and MillionVerifier. Second, a live-send bounce count on a warmed subset, the instrument Dropcontact uses, which is stricter than any SMTP check. Third, publishing the raw file, which is constrained by the fact that it is personal data. Fourth, a larger sample, since 500 leads carries a wider confidence interval than 5,000 or 20,000. Until those land, weight this the way you weight any first-party test in this category, ours included.
“Single-source coverage looks fine on a demo of 20 rows. Run 500 real leads through it and the gaps are impossible to hide. That is why we test on a real list, not a curated one.”
Has Cleanlist run more than one enrichment test?
Yes. Cleanlist has published three separate tests at three sample sizes, and the 500-lead benchmark is the largest of them. The 100-lead test ran the same list through the Cleanlist waterfall and through Apollo as a single-source control, returning 98% verified email and 85% phone for the waterfall against roughly 80% and 45% for Apollo. The 500-Lead Enrichment Benchmark, 2026 widened the control from one vendor to a range across single-source databases, returning 98% verified email and 85% direct dial against 70-80% email and 30-60% phone. The third test, run in March 2026 on 1,000 leads, contains no Cleanlist arm at all: it measured Apollo against ZoomInfo directly.
Why do the 100-lead and 500-lead tests report the same 98% and 85%?
Because they measured the same waterfall and it performed the same way twice. The figures are not copied between writeups. What changed between the two tests is the control, not the Cleanlist arm: the 100-lead test used Apollo alone as the comparison, and the 500-lead benchmark replaced that single vendor with a range across several single-source databases so no one vendor carried the result. A skeptical reader should treat the repetition as weak evidence on its own, since both tests are first-party and share a design. The honest reading is that two runs at different sizes did not contradict each other, which is a lower bar than independent replication.
What did the 1,000-lead March 2026 test measure?
It compared Apollo and ZoomInfo against each other on 1,000 leads processed on the same day, with no Cleanlist arm in the test. ZoomInfo returned a verified email for 84% of leads against Apollo's 78%, a 6-point gap. The wider gap was phone: ZoomInfo returned a mobile number for 67% of leads against Apollo's 41%, 26 percentage points. Job title accuracy was effectively tied, both vendors above 89%, which makes it a poor differentiator despite how often it appears in sales decks. The full methodology and result tables are published in Apollo vs ZoomInfo. Cleanlist ran the test, so it carries the same first-party caveat as everything else on this page.
How accurate is B2B contact data according to third-party benchmarks?
Published third-party benchmarks put B2B email finder coverage between 41.3% and 87.1% of submitted contacts. Anymail Finder's June 19 2026 test of 5,000 contacts found coverage from 41.3% (Voila Norbert) to 87.1% (FullEnrich), with a median across its 14 tools of 58.9%. Dropcontact's 20,000-contact test, updated February 2026, found real enrichment rates from 14.2% (FindThatLead) to 54.9% (Dropcontact). Prospeo's roundup states that "independent benchmarks show real enrichment rates ranging from 17% to 77%," which it calls "far below the 95%+ that most tools claim in marketing." No published test in this category has produced a number above 90% on submitted contacts.
What did Anymail Finder's 2026 benchmark measure?
Anymail Finder tested 14 email finders against 5,000 B2B decision-maker contacts sourced from LinkedIn Sales Navigator on June 19, 2026. The sample split roughly 2,570 United States, 810 United Kingdom, 810 France and 810 Germany. Every returned email was checked by three independent third-party verifiers, BounceBan, ZeroBounce and MillionVerifier, and counted valid only on a 2-of-3 consensus. On catch-all domains, BounceBan adjudicated. The benchmark data is published under a CC BY 4.0 licence, which is the strongest transparency posture in this category. Anymail Finder placed itself second on both coverage and validity rather than first.
| Tool | Coverage (of 5,000 submitted) | False positives | Validity of returned emails |
|---|---|---|---|
| FullEnrich (waterfall) | 87.1% | 4.0% | 95.7% |
| Anymail Finder | 86.4% | 0.9% | 98.9% |
| GetProspect | 71.2% | 3.0% | 96.0% |
| Findymail | 70.9% | 3.3% | 95.6% |
| Dropcontact | 69.4% | 5.1% | 93.1% |
| Apollo | 68.1% | 6.5% | 91.3% |
| Clearout | 60.2% | 15.1% | 79.9% |
| Hunter | 57.6% | 9.3% | 86.1% |
| Skrapp | 49.4% | 17.5% | 73.8% |
| Icypeas | 49.0% | 0.5% | 99.1% |
| Snov.io | 46.1% | 15.2% | 75.3% |
| FindThatLead | 45.7% | 23.2% | 66.3% |
| Prospeo | 45.2% | 3.7% | 92.5% |
| Voila Norbert | 41.3% | 13.5% | 75.4% |
Source: Anymail Finder, Email Finder Benchmark, tested June 19 2026.
What did Dropcontact's 20,000-contact benchmark measure?
Dropcontact submitted one identical file of 20,000 real contacts to 15 tools and measured results by sending live email and counting actual hard bounces, with a dual-entry manual domain audit on top. The sample split 9,800 United States, 9,700 European and 500 rest of world, and the study was updated in February 2026. Its headline metric is the real enrichment rate, defined as raw rate minus hard bounces minus wrong domains, which is a stricter test than verifier consensus because it catches addresses that pass SMTP checks and still fail on delivery. Dropcontact withholds the raw 20,000-contact file, stating that publishing personally identifiable data "would constitute processing without a legal basis." The full protocol is published. Dropcontact ranks itself first.
| Tool | Real enrichment rate | Hard bounce rate | Wrong domain rate | Total error rate |
|---|---|---|---|---|
| Dropcontact | 54.9% | 0.9% | 1.0% | 1.9% |
| GetEmail | 48.4% | 4.6% | 22.5% | 27.1% |
| FullEnrich | 48.3% | 3.6% | 11.7% | 15.3% |
| AnyMail Finder | 41.3% | 15.8% | 9.5% | 25.4% |
| Enrow | 40.9% | 2.3% | 5.8% | 8.1% |
| Findymail | 39.9% | 1.1% | 5.2% | 6.2% |
| Datagma | 37.4% | 10.6% | 12.8% | 23.5% |
| BetterContact | 37.2% | 4.7% | 6.4% | 11.2% |
| Aeroleads | 35.5% | 15.0% | 14.7% | 29.7% |
| Hunter | 32.5% | 11.2% | 5.3% | 16.4% |
| Icypeas | 31.6% | 1.0% | 5.8% | 6.8% |
| GetProspect | 26.1% | 8.3% | 9.5% | 17.8% |
| VoilaNorbert | 24.2% | 13.7% | 8.8% | 22.5% |
| LeadMagic | 22.6% | 10.6% | 5.4% | 15.9% |
| FindThatLead | 14.2% | 15.1% | 7.6% | 22.7% |
Source: Dropcontact, Email Finder Benchmark, updated February 2026.
Why do the Anymail Finder and Dropcontact benchmarks disagree?
They disagree because each publisher's own tool wins its own test and because they use different measuring instruments. Anymail Finder's test scores Anymail Finder at 86.4% coverage and Dropcontact at 69.4%. Dropcontact's test scores Dropcontact at 54.9% real rate and AnyMail Finder at 41.3% with a 15.8% hard bounce rate. The instrument explains most of the rest: verifier consensus on SMTP-style checks is more permissive than a live send, so any study measuring real bounces produces systematically lower numbers than one measuring verifier agreement. Sample geography does not explain the gap, since both studies are close to half European.
Same tools, two published tests, two very different answers
| Category | Anymail Finder study (coverage) | Dropcontact study (real rate) |
|---|---|---|
| Anymail Finder | 86.4% | 41.3% |
| Dropcontact | 69.4% | 54.9% |
| Findymail | 70.9% | 39.9% |
| Hunter | 57.6% | 32.5% |
What did Clay's work email provider test measure?
Clay tested 12 work-email providers, Dropcontact, Findymail, Hunter, LeadMagic, Nimbler, People Data Labs, RocketReach, Wiza, Icy Peas, Datagma, Prospeo and Snov, across two segments: 1,075 contacts at companies under 500 employees and 3,719 contacts at companies with 500 or more. Clay publishes no aggregate coverage percentages and no provider ranking at all. It publishes a confidence-scoring method instead: deliverability confidence of 100% if the contact replied, 80% if the address passed Instantly's verification, 0% if confirmed undeliverable, plus a separate correct-person score of 100%, 80%, 60% or 0%. Clay states that a provider's "overall confidence scores can drop as low as -25% when they deliver incorrect emails," because it treats a wrong address as more harmful than a missing one.
What is the difference between find rate and validity rate?
Find rate and validity rate are two different measurements that vendors routinely present as one number. Find rate, also called coverage or match rate, is the share of the contacts you submitted for which the tool returned an email. Validity rate, also called accuracy or deliverability, is the share of the emails it returned that passed verification. The denominators differ: find rate divides by what you sent, validity divides by what came back. A tool that returns an email for 45 of 100 submitted contacts and gets 41 of those right can honestly advertise "92% accuracy" while covering under half your list. Cleanlist's 98% is a coverage figure, measured after verification, on the 500 leads submitted.
Which vendor best shows the gap between coverage and accuracy?
Prospeo is the clearest worked example in the published data. In Anymail Finder's benchmark, Prospeo returned an email for 45.2% of the 5,000 submitted contacts, and 92.5% of the emails it returned were valid. On Prospeo's own site, the headline figure is "98% email accuracy." Icypeas shows the same shape from the other direction: 49.0% coverage, the lowest false-positive rate in the test at 0.5%, and the highest validity score at 99.1%. A buyer reading only the validity column would rank Icypeas first and FullEnrich fourth. A buyer reading only the coverage column would rank FullEnrich first and Icypeas tenth.
Both figures are defensible. Coverage divides by contacts submitted. Accuracy divides by emails returned. They are not comparable and are frequently quoted as if they were.
Source: Anymail Finder Email Finder Benchmark, June 19 2026; Prospeo published claimHow does Cleanlist's 98% compare with the 87.1% ceiling in published tests?
The two numbers come from different inputs and different architectures, and the comparison is not apples to apples. Four differences account for the distance. Cleanlist's run cascaded 15+ providers per lead; most tools in the third-party tests are single-source, and the one waterfall product measured, FullEnrich, took the top coverage spot at 87.1%. Cleanlist's 500 leads were weighted toward North America; both third-party samples are close to half European, and European coverage is structurally thinner. Cleanlist's input carried a LinkedIn URL per lead, which resolves identity more reliably than name plus domain. And 500 leads is a smaller sample than 5,000 or 20,000, so the confidence interval is wider. Run the protocol below if you want your own answer.
Does querying more providers make the data worse?
It can, and the published evidence shows exactly where. In Anymail Finder's test the waterfall product FullEnrich led on coverage at 87.1%, but its false-positive rate of 4.0% was more than four times Anymail Finder's 0.9%. In Dropcontact's stricter test FullEnrich reached 48.3% real rate, third overall, with an 11.7% wrong-domain rate. Querying more sources raises the chance that somebody holds the record, and it also raises the chance of importing a bad one. The only thing that prevents the second effect is verification sitting after the cascade rather than before it, which is why the Cleanlist run counted a lead as covered only after the returned email passed a deliverability check and the returned number passed a connectivity check.
What do the vendors themselves claim, and does it hold up?
Vendor marketing claims cluster at 95% to 98%, roughly 36 points above the 58.9% median coverage of the 14 tools in Anymail Finder's test, and most of that distance is denominator rather than dishonesty. Verified on August 2, 2026: Lusha states "98% email accuracy and 85% US, 86% global phone accuracy, refreshed daily against 300M+ verified contacts." UpLead states "95% data accuracy" and a "95%+ accuracy guarantee," backed by "Friendly Credits" that refund bounces. Findymail guarantees "under 5% bounce rate, or we refund your credits." Cognism claims "up to 3x higher connect rates" and publishes no accuracy percentage. Apollo claims "best-in-class email accuracy" with no figure, plus "under 1% invalid phone numbers." Hunter publishes a confidence score and no percentage.
Has ZoomInfo or Cognism been included in any published email accuracy benchmark?
No. Neither ZoomInfo nor Cognism appears in any of the three published third-party tests reconciled on this page. Anymail Finder's 14-tool test, Dropcontact's 15-tool test and Clay's 12-provider test all cover email finders and enrichment APIs, and none of them submitted a list to ZoomInfo or Cognism. Apollo is the only large sales-intelligence platform measured in any of them, at 68.1% coverage and 91.3% validity in Anymail Finder's June 2026 test. Cognism publishes no accuracy percentage on its homepage, claiming "up to 3x higher connect rates" from phone-verified mobiles instead. If you are comparing ZoomInfo or Cognism on email accuracy, there is no published measurement to compare them on, so run your own list through both.
Which vendor claims could not be verified for this page?
Four claims commonly repeated about this category could not be verified on August 2, 2026, and are excluded rather than repeated. Amplemarket's frequently quoted "under 3% bounce rate" does not appear on its homepage; what appears there is a relative claim that its data is "trusted by industry leaders to reduce bounce rates by 50%." Instantly's "80%+ inbox placement" does not appear on its homepage, which names inbox placement as a feature without a percentage. Ocean.io's 82% CRM match rate could not be checked because ocean.io returned HTTP 403 to an automated fetch. Prospeo's "92% API match rate" could not be located on a fetchable Prospeo page, though its 98% accuracy claim is published. If you see those four numbers cited anywhere, ask for the source page.
How do catch-all domains distort email accuracy figures?
Catch-all domains accept any address at the domain, so a standard SMTP check cannot tell a real mailbox from an invented one, and every study has to make a judgment call. Anymail Finder handles this by having BounceBan adjudicate catch-all results rather than treating them as valid by default. Dropcontact sidesteps the question by sending real email and counting hard bounces, which resolves catch-alls empirically. The Cleanlist 500-lead run flagged catch-alls and re-verified them rather than counting them as covered by default. When you compare two vendors' accuracy claims, the catch-all rule is usually the single largest hidden difference between them, and it can move a headline number by double digits on its own.
What is the difference between a hard bounce and a soft bounce?
A hard bounce is a permanent delivery failure and a soft bounce is a temporary one, and only hard bounces are usable as an accuracy measurement. A hard bounce means the receiving server rejected the address outright, usually because the mailbox does not exist, and it will fail the same way on every retry. A soft bounce means the message was refused for a transient reason such as a full mailbox, a size limit, or a server that was briefly unavailable, and the same address may deliver later. Dropcontact's benchmark counts only hard bounces in its real enrichment rate, and that column across its 15 tools runs from 0.9% (Dropcontact) to 15.8% (AnyMail Finder). Discard soft bounces rather than scoring them as bad data.
How should you score a published data accuracy benchmark?
Score a published benchmark on five questions before you quote it, including this one. First, sample size and composition: how many contacts, from where, at what seniority and company size. Second, is the verifier independent and named, or in-house. Third, is the raw data published, and under what licence. Fourth, does it measure real hard bounces on a live send, or only verifier agreement. Fifth, did the publisher rank itself first. A study can fail the fifth test and still be worth citing if it passes the first four, which is the situation across this entire category, ours included.
| Study | Sample | Independent verifier | Raw data published | Publisher ranked itself first |
|---|---|---|---|---|
| Cleanlist 500-Lead Benchmark, 2026 | 500 leads, stratified by industry, seniority and size, NA-weighted | No, in-house verification and connectivity checks | No, personal data withheld; protocol published in full | Yes |
| Anymail Finder, Jun 19 2026 | 5,000 contacts, US/UK/FR/DE | Yes: BounceBan, ZeroBounce, MillionVerifier, 2-of-3 consensus | Yes, CC BY 4.0 | No, placed itself 2nd on both axes |
| Dropcontact, updated Feb 2026 | 20,000 contacts: 9,800 US, 9,700 European, 500 rest of world | No, in-house live send plus dual-entry domain audit | No, withheld on GDPR grounds; protocol published | Yes |
| Clay, no date stated | 1,075 SMB plus 3,719 enterprise contacts | No, in-house confidence scoring | No | Not applicable, Clay publishes no ranking |
| Prospeo roundup | No original test, aggregates others | Not applicable | Not applicable | Yes, quotes itself at 98% |
How can you run this protocol against Cleanlist or anyone else?
You can run a defensible version of this test in an afternoon, and it works against us as well as against anyone we compete with. Pull a random sample of 300 to 1,000 real contacts from your CRM or target list, stratified by country, seniority and company size, and do not hand-pick easy rows. Strip each record to first name, last name, company name and either company domain or LinkedIn URL, removing any existing email or phone so you are measuring the vendor and not your old data. Submit the identical file to every vendor on the same day. Then measure three numbers separately rather than one blended number, and publish which one you are quoting.
- Find rate = emails returned divided by contacts submitted.
- Validity rate = emails that pass verification divided by emails returned. Use at least two verifiers you pay for yourself, count a result valid only on agreement, and record catch-alls in their own bucket rather than folding them into either column.
- Real rate = emails that actually deliver divided by contacts submitted. Send a warmed subset, wait 72 hours, subtract hard bounces and wrong-domain results. This is the number that predicts your outbound outcome.
The denominator is the part that trips people up, and getting it wrong moves the same measurement by 50 points. Divide by contacts submitted and you get coverage, which is what you actually care about, because the rows a vendor skipped are rows your reps cannot work. Divide by emails returned and you get validity, which tells you how much of what came back is safe to send to. Prospeo's 45.2% coverage and 92.5% validity in Anymail Finder's test are the same tool in the same run. Cleanlist's 98% divides by the 500 leads submitted. Never let a vendor blend the two, and never blend them yourself.
For phone, run the same shape: dial-connect rate divided by contacts submitted, counting only numbers that reach the named person. Report your sample composition and your catch-all rule alongside the three numbers. If you do that and publish it, you will have produced better evidence than most of what is currently cited in this category.
What should you ask a data vendor before you buy?
Ask five questions and refuse blended answers. First, is your headline percentage a find rate or a validity rate, and what is the denominator. Second, who verified it, and can I see their name. Third, what was the sample: how many contacts, which countries, which company sizes. Fourth, how do you treat catch-all domains in that number. Fifth, will you run my list, unmodified, and let me verify the output with a verifier I choose. A vendor that will not answer the fifth question is telling you something about the first four. Ask us the same five, and the answers are on this page.
“Every accuracy number in this category is answering one of two different questions, and almost nobody says which. Find rate asks how many of the contacts you submitted came back with an email. Validity rate asks how many of the emails that came back were real. A tool can score 45% on the first and 92% on the second, and both numbers go in the same marketing sentence.”
Run this benchmark on your own list
Upload a sample and measure your real coverage rate before you commit to any vendor. The free tier is 30 credits a month with no card. Enrichment costs 1 credit per verified email and 10 per phone, and people search is free.
Frequently asked questions
What is the Cleanlist 500-Lead Enrichment Benchmark?
It is a first-party test in which Cleanlist ran 500 real B2B leads, stratified by industry, seniority and company size and weighted toward North America, through a 15+ provider waterfall and separately through single-source databases on identical input. The waterfall returned a verified email for 98% of leads and a direct-dial or mobile number for 85%. Single-source returned 70-80% verified email and 30-60% phone on the same leads. Coverage counted only after verification. Cleanlist designed, ran and scored the test, and the protocol is published above so anyone can repeat it.
Was the Cleanlist benchmark verified by a third party?
No. Cleanlist ran and scored it in-house, which is also true of the Anymail Finder, Dropcontact and Clay tests reconciled on this page. What we publish instead of an independent audit is the protocol in full: sample construction, input format, what counted as a hit, how catch-alls were handled, and why single-source is reported as a range. Independent verification of the returned records by verifiers we do not pay would make it stronger, and it has not happened yet.
Why does Cleanlist report 98% when third-party tests top out at 87.1%?
Because the inputs and architectures differ. The 87.1% ceiling belongs to FullEnrich in Anymail Finder's test, the only waterfall product measured there, on a sample that is roughly half European with name-plus-domain input. Cleanlist's 500 leads were North America weighted, carried a LinkedIn URL per lead, and were cascaded across 15+ providers with verification after the cascade. Sample size differs too, 500 versus 5,000. Both figures are coverage of submitted contacts, but they are not measuring the same test.
What is the most accurate B2B email finder in 2026?
It depends on which metric you mean, and the published tests disagree. On validity of returned emails, Anymail Finder's June 2026 benchmark ranks Icypeas first at 99.1%, then Anymail Finder at 98.9%, then GetProspect at 96.0%. On coverage of submitted contacts, the same test ranks FullEnrich first at 87.1%, then Anymail Finder at 86.4%. On real enrichment rate net of hard bounces, Dropcontact's 20,000-contact test ranks Dropcontact first at 54.9%. Each publisher's own product performs best in its own study, so read all three rankings and weight them accordingly.
Why do vendors claim 98% accuracy when benchmarks show 41-87%?
Because the two figures use different denominators and are both technically defensible. A 98% accuracy claim usually means 98% of the emails the tool returned passed verification. A 41-87% benchmark figure means that share of the contacts you submitted got an email back at all. A tool that returns emails for 45% of your list and gets 92.5% of those right, which is Prospeo's measured result in Anymail Finder's test, can publish "98% accuracy" and still leave more than half your list unenriched. Always ask which denominator you are being quoted.
Which published benchmark has the best methodology?
Anymail Finder's June 2026 benchmark has the strongest methodology on transparency: 5,000 contacts across four countries, three named independent verifiers under a 2-of-3 consensus rule, explicit catch-all adjudication, and data published under CC BY 4.0. Dropcontact's has the strongest measurement instrument, since counting real hard bounces on a live send of 20,000 contacts catches failures that SMTP verification misses, though it withholds raw data on GDPR grounds and ranks itself first. Reading both, alongside our 500-lead protocol, is better than trusting any one of them alone.
What does "real enrichment rate" mean?
Real enrichment rate is Dropcontact's metric, defined as raw find rate minus hard bounces minus wrong domains, measured by sending live email to every returned address. It is the strictest published measure in this category because it counts only addresses that survived actual delivery, divided by the contacts you submitted. Across Dropcontact's 15-tool test it ranges from 14.2% to 54.9%. Prospeo summarises the published range across studies as 17% to 77%.
Can I cite the Cleanlist 500-lead benchmark?
Yes. Cite it as "Cleanlist 500-Lead Enrichment Benchmark, 2026," note that it is a first-party test, and link back to this page. The protocol above is the full method, so a reader can check what was measured and repeat it. The raw list is not published because it is personal data.
Sources on this page
References & Sources
- [1]
- [2]
- [3]
- [4]
- [5]
- [6]
- [7]
- [8]
- [9]
- [10]
- [11]
- [12]
- [13]
For related reading on this site, see the best data enrichment tools for 2026, the State of B2B Data Quality 2026, the best B2B data providers, and how waterfall enrichment and email verification work.
Run this on your own list
Upload a CSV, enrich every row across 15+ providers, and export the result back to your CRM. First 30 rows free.
Enrich 30 rows free30 credits free every month · No credit card

