Email A/B testing: a practical guide

TL;DR, the essentials
- A/B testing (or split testing) compares two versions of an email on a sample, then sends the winner to the rest of the list.
- You can test the subject line, the sender name, the content, the call-to-action (CTA) button and the send time.
- The golden rule: one variable at a time, otherwise you cannot tell what made the difference.
- A result is only trustworthy if it is statistically significant (aim for 95% confidence) on a large enough sample.
Torn between two subject lines, two buttons or two send times? Instead of guessing, email A/B testing lets your recipients settle it for you, with numbers to back it up. It is the simplest way to stop steering on instinct and improve campaign after campaign. Here is what you can test, how to run it without tripping yourself up, and above all how to read the results without kidding yourself.
What is A/B testing in email marketing?
A/B testing, also called a split test, means creating two versions of the same email that differ on a single element (variant A and variant B). You send each version to a small sample of your list, measure which one performs better, then roll out the winner to all the remaining contacts. In practice, the process comes down to three steps.
Isolate one variable
Subject line, sender, CTA… You change one thing between A and B. Everything else in the email stays strictly identical.
Test on a sample
The tool sends A and B to two randomly drawn subgroups, for example 10 to 20% of the list split into two equal halves.
Send the winner
After a set delay, the better-performing version goes out automatically to the rest of the database.
Most email marketing platforms build this feature in and automate the whole flow, including picking the winner and sending the final version. A/B testing is not reserved for experts: it is a reflex you can adopt from your very first campaign.
Keep in mind
A test only holds up if A and B are identical in every way except the element being tested. Change two things at once, and the result no longer means anything.

What can you test in an email?
Almost everything can be tested, but some elements carry far more weight than others. Here are the five classic levers, from the most rewarding to the most refined:
- The subject line: this is the most common test and often the most profitable, because it decides the open. Compare short against long wording, a question against a statement, with or without an emoji, with or without the first name.
- The sender name: “Marie from MyBrand” against “MyBrand”, or a human address against a
noreply@. The impact on trust and opens can be clear. - The content: text length, images or no images, block order, tone of voice. These tests mainly move engagement and clicks.
- The call-to-action (CTA) button: its wording (“Grab the deal” against “See the offer”), its color, size or position. This is a direct lever on the click rate.
- The send time and day: Tuesday 10 a.m. against Thursday 2 p.m., for example. Careful though, this test follows special rules we detail further down.
A word on priority: tests on the subject line and the CTA generally deliver more value than tests on the sender name or the footer. Start there. Note too that the subject line mainly influences the open rate, while the content and CTA act on the click rate.
Not every tool tests the same way
Some platforms limit A/B testing to the subject line, others test the full content. Our comparison sorts out which is which.
What method should you follow for a reliable test?
A badly framed A/B test produces false conclusions, which is worse than no test at all. Three rules keep the approach solid.
One variable at a time. This is the non-negotiable rule. If you change both the subject line and the CTA, and version B wins, you will never know which of the two made the difference. To test several elements, run several successive tests, or use a dedicated multivariate test if your list is very large.
A clear hypothesis before you launch. Do not test “just to see”. Frame an intention: “I believe a subject line with the recipient’s first name improves opens”. That way you know what to measure and what to do with the result.
The right success metric. Choose the metric before the test, based on what you are testing. A subject line test is judged on the open rate, a CTA test on the click rate, an offer test on conversions. Judging a subject line test on clicks makes no sense.
Good habit
Since Apple Mail preloading, the open rate is partly inflated. For a subject line test, keep the open as a reference, but confirm with the click rate whenever you can.
Quick quiz
How many variables should differ between version A and version B?
What sample size and duration do you need?
This is the point that separates a real test from a lucky guess. A gap measured on 40 recipients proves nothing; on several thousand, it becomes credible.
Sample size. Common benchmarks point to a minimum of around 1,000 recipients per variant to hope for a reliable result, so at least 2,000 contacts involved in the test. It is not a magic threshold: the smaller the performance gap you expect, the more volume you need to detect it. Below a few hundred contacts per version, an A/B test has little chance of settling anything clearly, and sometimes you are better off simply sending your best hunch.
Indicative benchmarks, July 2026, aggregated from provider guides and email marketing platforms. Exact thresholds depend on your baseline open rate and the size of the difference you are chasing.
Duration. Let the test run long enough to smooth out swings from one moment of the day to another. A window of 3 to 7 days is often recommended for standard campaigns; some tools pick a winner after just a few hours, which can be enough for a simple subject line test but stays more fragile. In B2B, stretch the window: your contacts check their email mostly during office hours.
The send-time test trap
Testing two send times is useful, but a given contact can only receive the email once. So this test compares two different groups, not to be confused with a subject line test on a single send. Repeat it across several campaigns before turning it into a rule.
How do you interpret A/B test results?
One version wins by two-tenths of a point? That is not necessarily a winner. The real question is whether the gap is statistically significant, meaning too pronounced to be down to chance. The convention in email marketing is to aim for a 95% confidence level: below that, the result is treated as inconclusive.
In practice, most platforms display a confidence indicator directly, or only declare a winner once the threshold is reached. If your tool does not, hold on to this simple principle: a small gap on a small sample proves nothing, a clear gap over several thousand sends is credible. Three reflexes to conclude well:
- Check significance before acting. As long as the 95% threshold is not met, treat the test as unsettled and do not turn it into a general rule.
- Tie the result to a real goal. A better open rate that produces neither clicks nor conversions brings nothing. Trace the chain up to the metric that actually matters to you.
- Document and repeat. Log the hypothesis, the winner and the gap. A/B testing is a cycle: each test feeds the next, campaign after campaign.
A test that leads to no decision, or whose gap is not significant, is not a failure: it is a hypothesis ruled out, and time saved.The MiisterSoftware team, the principle of continuous optimization.
A/B testing depends mostly on your platform
Statistical confidence, automatic winner send, full-content testing: see which tools really do it.
The next step
Ready to test your campaigns? Compare the platforms in our guide to the best email marketing software 2026. To go further on the metrics behind your tests and how to read your own reports, keep exploring the email marketing hub.
Frequently asked questions
What is A/B testing in email marketing?
It is a method that compares two versions of an email (A and B), differing on a single element, by sending them to two samples of the list. You measure which one performs better, then send the winning version to the remaining contacts. The goal is to decide with data rather than on instinct.
What can you test in an email?
The five most tested elements are the subject line, the sender name, the content (text, images, layout), the call-to-action button and the send time or day. Tests on the subject line and the CTA are generally the most rewarding.
How many contacts do you need for a reliable A/B test?
Common benchmarks point to at least 1,000 recipients per variant, so about 2,000 contacts involved, to hope for a credible result. The exact threshold depends on your baseline open rate and the size of the gap you are chasing. Below a few hundred per version, a test rarely settles anything.
How long should an A/B test run?
A window of 3 to 7 days is often recommended to smooth out variations across the day. Some tools declare a winner after a few hours, which can be enough for a subject line test but stays more fragile. In B2B, extend the duration because emails are read mostly during office hours.
How do you know if the result is reliable?
The result must be statistically significant, meaning too pronounced to be down to chance. The convention in email marketing is to aim for a 95% confidence level. Most platforms display this indicator. A small gap on a small sample proves nothing.
Can you test several elements at once?
Not in a standard A/B test: the rule is to change one variable only, otherwise you cannot tell what made the difference. To test several elements at once, you need a multivariate test, which is reserved for very large lists because it demands far more volume.