Measurement and iteration — Reply rate as the survivor metric, the sevenfold benchmark spread, and why small A/B tests mislead
Which numbers mean something, what "good" looks like this year, and how to find out why a campaign underperformed. Three disciplines run through the whole module.
Benchmarks are always dated, and always divided by something. Published 2026 averages range from 0.45% to 3.43% reply rate — a sevenfold spread that is almost entirely definitional. A number without its year, its corpus and its denominator is not evidence.
Open rate is broken. Mail clients pre-load tracking images, so a large share of recorded opens have no human behind them. Replies are the metric that survived.
Most cold-email A/B tests cannot detect what they claim to. At realistic reply rates, telling a good subject line from a bad one needs thousands of prospects per variant, not the hundred that common guidance suggests. Knowing that changes what is worth testing.
The pages:
- Core metrics — Reply rate as the one metric that survives scrutiny, the health-check thresholds around it, and why the same campaign can honestly read 0.45% or 3.4%.
- Benchmarks — The 2026 published figures with owners and dates, why 0.45% and 3.43% are both honest averages, and the case for benchmarking against your own last campaign.
- Diagnosing underperformance — The delivery → list → offer → copy diagnosis order, a symptom-to-layer table with the out-of-office tell for silent placement failure, and why copy comes last.
- A/B testing — Why the published 100-per-variant guidance detects only implausible effects at a 3.4% baseline — the real number is near 2,200 per arm — and what to do instead.
- Tracking caveats — How Apple's 2021 Mail Privacy Protection broke the open pixel, the tracking-on versus tracking-off schools, and the default of tracking off for cold campaigns.
- Scaling what works — Confirming a result is real, the four-step scaling order, the depth effect that makes the second thousand contacts worse, and the month of sender lead time.
Instrumentation is a precondition, not an afterthought: a campaign recorded only as an average cannot be diagnosed, only re-run.
- Up: Playbook contents
- Previous: Replies and objections