Email a/B Testing: How to Run Tests That Lift Revenue
Email A/B testing is the practice of sending two versions of a campaign to small audience segments, measuring which performs better, then sending the winning version to the rest of your list. It turns guesswork into evidence, so your open rates, click rates, and revenue improve with every send instead of drifting on opinion.
Why Email a/B Testing Beats Guessing
Every list is different. The subject line that doubled another brand's opens can fall flat for yours because the audience, offer, and timing are different. A/B testing removes the debate by letting subscribers decide. Instead of the highest-paid opinion winning, the version that earns the click wins, and you bank that learning for future campaigns.
The payoff compounds. A small, repeatable win on open rate or click rate, applied across dozens of sends per quarter, adds up to a materially larger pipeline. Testing also protects you from drift: as your list grows and ages, what worked six months ago may no longer work, and only live tests catch that.
What You Can Test in an Email
Almost any element of a message can be varied. The highest-leverage tests, in rough order of impact, are:
- Subject line - the single biggest driver of opens. Test length, tone, personalization, questions, urgency, and emojis.
- Preview text - the snippet after the subject in the inbox. It can reinforce or undermine the subject.
- From name - a person's name versus a brand name, or a person plus the company.
- Send time and day - when your segment actually opens mail.
- CTA placement and wording - button text, color, and whether the ask appears once or twice.
- Content structure - short versus long, one offer versus multiple, image-heavy versus text-first.
Early on, prioritize subject lines and send time because they are cheap to change and affect every send. Save layout and content tests for when you have a stable baseline.
How to Run an Email a/B Test Step by Step
Follow a consistent process so results are comparable over time:
- Pick one variable. Test a single element, such as the subject line, so you know what caused any difference. Testing two changes at once leaves the result ambiguous.
- Form a hypothesis. State what you expect and why, for example, "A question subject line will lift opens because it triggers curiosity." The hypothesis is what makes the test reusable.
- Choose the split. Send each version to a random 10 to 20 percent of your audience, held back from the rest until the test closes.
- Set a sample size and deadline. Decide in advance how many recipients and how many hours the test needs before you read it.
- Run it, then roll out the winner. Send the better version to the remaining holdout as soon as the test reaches significance.
- Record the learning. Note the winner, the lift, and the conditions, so the next campaign starts from evidence rather than a blank page.
How Big Should Your Sample Size Be
Small lists need a larger share of the audience in the test to reach a confident result. The table shows how much of your list each variant needs for a typical medium effect, assuming a roughly even split between the two versions.
| Total list size | Recipients per variant | Share of list in test |
|---|---|---|
| 1,000 | ~250 | 50 percent |
| 5,000 | ~300 | 12 percent |
| 20,000 | ~350 | 3.5 percent |
| 100,000 | ~400 | 0.8 percent |
Above about 20,000 subscribers you can usually test on a small fraction of the list and still get a clear read, which means testing rarely costs you meaningful reach.
How Long Should an Email a/B Test Run
Most email tests should close within 4 to 24 hours. Email engagement is front-loaded: the large majority of opens happen in the first few hours after send. Running longer mostly adds stale recipients who behave differently and muddies the comparison. For time-of-day tests, run the variants on different days rather than holding one open for days.
Common Mistakes That Ruin Email Tests
- Testing too many things at once. If you change the subject and the CTA, you cannot tell which one moved the number.
- Reading results too early. Calling a winner before the sample is large enough produces false wins that reverse at scale.
- Unequal or non-random splits. If one variant goes to your most engaged segment by accident, the test is invalid.
- Testing on tiny lists. With a few hundred subscribers, random noise swamps any real difference; aggregate learning over several sends instead.
- Ignoring the holdout. If you never send the winner to the rest, you captured insight but left revenue on the table.
How to Read the Results and Act on Them
Look at the primary metric you chose before the test, usually open rate for subject tests or click rate for content and CTA tests. Check whether the difference is statistically significant, not just numerically larger. Most email platforms report this, or you can use a simple two-proportion test. If the result is significant, adopt the winner and write down the rule. If it is not, treat it as inconclusive and fold the attempt into a larger pooled analysis rather than over-fitting to noise.
Build a lightweight test log: variable, hypothesis, result, and lift. Over a year this becomes your team's private playbook of what actually works for your audience. Pair disciplined testing with strong deliverability and segmentation, and the same list will keep producing more revenue without growing headcount. For the measurement side, our blended ROAS guide shows how to value the revenue those emails drive, and our guide to disapproved ads covers the compliance habits that keep acquisition channels healthy. For the broader program, see our B2B email marketing guide.
When Not to a/B Test
Testing is not always the right move. If your list is under a few hundred recipients, random noise will drown any real signal, so aggregate learning across sends instead of calling premature winners. If you are sending a one-time, high-stakes message such as a major product launch or a crisis note, do not gamble the result on a test; send your best version to everyone. And if you have no clear hypothesis, a test will produce a number without a lesson. Only test when you can act on the answer and learn something reusable.
How to Prioritize Your Testing Roadmap
You will never run out of things to test, so rank them by expected impact and cost. Start with the changes that touch every send and require no new design work: subject lines, preview text, and send time. Next come CTA and offer changes, which need more setup but move revenue directly. Save structural and visual redesigns for when you have a baseline strong enough to measure a real lift against. Keep a single running list of candidates, and pull the next one as soon as the current test closes, so testing becomes a habit rather than a quarterly scramble.
Tools and Built-In Testing in Email Platforms
Most modern email service providers include native A/B testing, often called split testing or experiments. They handle the random split, the holdout, and the automatic send of the winner, which removes the manual errors that sink do-it-yourself tests. Some platforms add predictive or machine-learning send-time optimization that essentially runs a continuous time test for you. Use the built-in tool for standard subject and content tests, and reserve custom setups for questions the platform does not support, such as cross-channel holdouts. Whatever you use, export the result into your test log so the learning survives a tool change.
Frequently Asked Questions
How Many People Do I Need to a/B Test an Email?
It depends on your list size and the size of the effect you expect. On a 1,000-person list you may need to test half of it to see a clear result, while a 100,000-person list can reach a confident read with under 1 percent in the test. Aim for at least a few hundred recipients per variant and use a significance check before declaring a winner.
How Long Should an Email a/B Test Run?
Most tests should run 4 to 24 hours. Email engagement concentrates in the hours right after send, so extending the window mainly adds late, low-intent opens that obscure the comparison. For send-time experiments, test different days rather than holding one variant open for days.
What Is the Best Thing to Test First in Email?
Start with the subject line. It is the cheapest change and the biggest lever on opens, which gates everything else, because no one clicks an email they never open. Once you have a stable open rate, move to preview text, send time, and then CTA wording.
Can I Test More Than Two Versions at Once?
Yes, most platforms support A/B/C/D tests, but each added variant shrinks the sample per version and slows the path to significance. For most teams a clean two-version test is clearer and faster. Use multivariate testing only when you have a large list and a specific question about how elements interact.
How Do I Know If My Email Test Result Is Real?
Check statistical significance, not just the raw difference. A two-percentage-point gap on 200 recipients can be noise, while the same gap on 10,000 is likely real. Use your platform's significance indicator or a two-proportion test, and require a pre-set sample size before you read the result.