Email Verification Accuracy: What 99% Really Means
In short
Headline accuracy figures are calculated on addresses that returned a clear answer, which excludes the catch-all segment where verification is genuinely hard. The metric worth comparing is what share of that ambiguous segment a tool resolves.
Open the pricing page of any email verifier and you will find a number in the high nineties. They are not competing on it so much as matching each other, which is itself informative: when every vendor in a category reports the same figure, the figure has usually stopped measuring anything that distinguishes them.
This is not an accusation of dishonesty. The numbers are generally true. The problem is what they are calculated over.
How the number gets built
To state an accuracy percentage you need a test set where the right answer is known in advance, then you compare the tool’s verdicts against it. The construction is sound. The question is which addresses make it into the test set.
Addresses with a known answer are, by definition, addresses whose mail servers gave a clear response. A server that says yes to one mailbox and no to another has told you the truth about both, and you can score a verifier against that. Addresses on domains that accept everything cannot be scored this way, because you have no independent way to establish ground truth without sending a message and watching what happens, which is the very thing verification exists to avoid.
So the hard segment is excluded from the denominator. Not through bad faith, but because including it is methodologically awkward. The result is an accuracy figure calculated exclusively on the addresses that were never in doubt.
A 99 percent score on the easy addresses is not a strong claim. It is close to a definitional one.
What accuracy actually measures
Accuracy in this context is a statement about error, not about coverage. It answers: of the verdicts this tool issued, how many were right? It says nothing about how many verdicts it declined to issue.
Those are separate properties and only one of them appears in the marketing. A tool can be extremely accurate and extremely unhelpful, by the simple method of refusing to answer whenever the answer is difficult. Every address it returns as unknown is an address that cannot contribute an error to its accuracy score.
There is an uncomfortable incentive buried in that. Under an accuracy-only metric, declining to answer is strictly safer than attempting a hard call, because a non-answer costs nothing while a wrong answer costs a percentage point. A metric that rewards silence will produce tools that are quiet.
One number, two very different errors
A single accuracy percentage also conceals which kind of mistake a tool makes, and the two are not equally expensive.
A false positive marks a dead address deliverable. You send, it bounces, and the cost lands on your sending domain. This is the error everyone thinks about, and it is at least visible: the bounce report tells you it happened.
A false negative marks a live address undeliverable. You delete a real contact you paid to source, and nothing anywhere tells you it occurred. There is no report for leads you quietly removed, no feedback loop, and no way to notice the pattern across campaigns. A tool that trims aggressively to protect its bounce numbers is invisibly expensive.
Both are counted identically in an accuracy figure, so two tools with the same score can behave very differently on your data. When you run the seed test described below, note which direction each tool errs in, because that tells you more than the aggregate.
Where the figure stops describing your list
On a typical B2B list, 25 to 30 percent of addresses sit on catch-all domains, which accept mail for every address they are asked about. That is the segment excluded from the accuracy calculation, and it is not a rounding error. It is a quarter to a third of the file, weighted toward larger organisations, which is to say weighted toward the contacts most people are paying to reach.
The consequence is that a 99 percent accuracy figure describes, roughly, 99 percent of the 70 to 75 percent of your list that was never difficult. What happens to the rest is not addressed by the number at all. If you are verifying a consumer list where the catch-all share is small, this barely matters. If you are verifying business contacts, it is most of what you wanted to know.
The metric that separates tools
The useful question is not how often a verifier is right. It is what proportion of the ambiguous segment it can answer at all. Call it resolution rate: of the addresses a first pass could not settle, how many end up with a definitive verdict?
This metric has three properties the accuracy figure lacks. It varies enormously between tools, so it actually discriminates. It describes the part of the list you are uncertain about, which is the part you were buying help with. And it cannot be improved by declining to answer, because declining to answer is precisely what it measures.
For reference, our second pass returns a definitive verdict on roughly 72 percent of the catch-all segment. That number is comparable in a way a headline accuracy figure is not, because a tool that does not run a second pass scores zero on it by construction rather than by performing badly.
Testing a verifier yourself
You do not need a vendor’s benchmark. A seed list you assemble yourself is more informative than any published figure, because it is built from addresses whose truth you already know.
- Collect known-good addresses. Ten or twenty mailboxes you control or have recently corresponded with. Any tool marking these undeliverable has failed in the most expensive direction, since false negatives silently delete real contacts.
- Collect known-bad addresses. Invent mailboxes at domains you know are not catch-all, ideally domains you control. You know for certain these do not exist, so anything reported as deliverable is a false positive.
- Establish a catch-all domain. Pick a domain and verify a random string at it. If that comes back deliverable, no such mailbox exists and the domain accepts everything. You now have a confirmed catch-all domain to sample real addresses from.
- Run all three groups and read the third one. The first two groups will separate careless tools from careful ones, and most credible verifiers pass both. The catch-all sample is where the differences live: count how many come back with a real verdict rather than a status.
A hundred addresses is enough to see the pattern, and the free checker handles the random-string step without an account. Run the same list through two tools and the comparison stops being about marketing copy.
Why a second pass changes the number
If the limiting factor is coverage rather than error, the way to improve is to go back to the unresolved addresses with different techniques rather than to refine the first attempt.
That is what two-pass verification does. The first pass behaves like any conventional verifier and sorts its leftovers honestly. The second runs only on that leftover segment, using deeper SMTP-level verification and waterfall routing to approach the mailbox by a route that behaves differently from the one that produced the ambiguity. Because it works on a minority of the list, it can afford to be slower and costlier per address than any first pass could.
The effect on a list is coverage, not correctness. The addresses that were already settled stay settled and the accuracy figure barely moves. What changes is how much of the file has an answer at all. The two-pass model covers the architecture, and catch-all domains covers why that segment behaves as it does.
Questions worth asking a vendor
None of these are hostile, and a vendor with a good answer will give it readily.
- What population is the accuracy figure measured over? Specifically, are catch-all addresses in the denominator? If they are excluded, the number describes the easy portion of a list.
- What share of a B2B list comes back unresolved? This is the coverage question, and it is the one that predicts how useful the output will be.
- What happens to that segment? If the answer is a status name, the tool has detected a condition rather than resolved it.
- Can I run my own seed list before committing? Any verifier confident in its results will let you test on a sample.
Ask us the same four. Our answers are that catch-all addresses are measured separately and reported as a resolution rate rather than folded into an accuracy claim, that a quarter of a typical B2B list arrives unresolved after a first pass, that a second pass resolves roughly 72 percent of it, and that a free account includes 100 credits specifically so you can check that on your own data. The verifier comparisons work through how each major tool handles the same segment.
Frequently asked questions
Is a 99% accuracy claim dishonest?
Usually not dishonest, but rarely as informative as it looks. The number is generally true for the population it was measured on, which is addresses that produced a clear server response. Since those are the addresses nobody was uncertain about, a high figure there is close to guaranteed. The claim becomes misleading only when it is presented as describing a whole list rather than the easy portion of one.
What accuracy should I actually expect?
On addresses that resolve cleanly, expect any credible verifier to agree with any other, and expect very few mistakes. On catch-all domains, expect a single-pass tool to return no verdict at all rather than a wrong one. The realistic question is not what percentage is correct but what percentage receives an answer, and that is where tools genuinely differ.
How can I test a verifier's accuracy myself?
Build a small seed list where you already know the answers: addresses you control, addresses you have confirmed are dead, and a sample from a domain you have established is catch-all. Run it through the tool and check three things. Did it catch the known-bad addresses, did it avoid marking known-good ones as invalid, and what did it do with the catch-all sample. The third answer is the one that varies.
Why do two verifiers disagree on the same address?
Almost always because the address sits in ambiguous territory rather than because one is wrong. Mail servers answer differently depending on who is asking, how often they have been asked, and what reputation the asking infrastructure carries. Disagreement clusters heavily on catch-all domains and greylisting servers, and it is rare on addresses that give a clean response.
Does higher accuracy mean lower bounce rates?
Only if the accuracy figure covers the addresses you actually send to. A tool that is 99 percent accurate on the addresses it resolves and silent on a quarter of your list leaves your bounce rate dependent on what you do with that quarter. Coverage and accuracy are separate properties, and coverage is the one usually left out of the marketing.