Back to the blog
e-mailmarketinga/b-testonderwerpregel
New

A/B testing subject lines: setup, sample size and pitfalls

An A/B test of subject lines works like this: you send the same mail with two subject lines to two random, equally sized parts of your list, measure which version produces more unique clicks, and send the winner to the rest. The result is only usable if the two versions differ in one thing, the groups are large enough to rule out chance, and you measure clicks rather than opens.

What you are actually testing

A subject line has one job: getting the mail opened. But because opens are no longer measured reliably since the privacy measures of mail programs, you measure the effect of the subject line best on what comes after: the number of unique recipients who click. A subject line that arouses a lot of curiosity but is not backed up by the content then rightly scores worse than an honest line that gets the right people to open. Why opens are unreliable is explained in our article on newsletter KPIs.

The setup in five steps

  1. Formulate a hypothesis. Not "which is better" but "a subject line with the concrete benefit does better than one with a question". Then you learn something you can use in following campaigns, even if the test turns out differently than expected.
  2. Change one thing. Length, or tone, or personalisation, or a number in it. Not everything at once, because then you will not know afterwards what made the difference.
  3. Split randomly. Your platform must shuffle the list and then split it. Giving group A the first half of an alphabetical list and group B the second seems harmless but is not: one domain or one sign-up period suddenly ends up in one group.
  4. Send both versions at the same moment. Version A at 9 and version B at 2 is a test of the send time, not of the subject line.
  5. Wait long enough and decide on unique clicks. A winner after one hour is a winner among people who read mail in the morning. Give the test at least the time your readers usually take to respond; for most newsletters that is several hours to a day.

How large must the sample be?

Larger than most people think. The difference you are trying to measure is usually small: a few per cent more clicks. To distinguish such a small difference from chance you need thousands of recipients per group, not hundreds. A practical rule of thumb: if both groups together count fewer than a few thousand addresses, the test is almost never reliable, and two consecutive campaigns asking the same question are often more useful than one split send.

If you have a large list, you do not need to put everything into the test. A common setup is 10 to 20 per cent of the list for the test (split across A and B) and the remaining 80 to 90 per cent for the winner. The smaller the test share, the longer you must wait and the larger the difference must be to conclude anything.

Five pitfalls

1. Measuring on opens

Already mentioned, but it is the most common mistake, because platforms set it that way by default. Set the winning metric explicitly to clicks.

2. Deciding too early

After an hour version A leads, after a day version B. Whoever lets the platform pick the winner automatically after an hour is measuring which subject line appeals to fast readers.

3. Treating the winner as proven truth

One test says something about one mail to one list on one day. Only after several tests pointing the same way may you derive a rule you apply to following campaigns.

4. Only looking at the short term

A sensational subject line wins the test and a week later produces more unsubscribes and complaints. After every test, also look at unsubscribes and complaints per group. A subject line that wins on clicks and loses on complaints is not a winner; it damages your deliverability over time.

5. Testing on a dirty list

A list with many dead or dormant addresses dilutes both groups equally, but makes the difference smaller and the test therefore less sensitive. Clean up first.

A practical example

A webshop with 30,000 active addresses tests two subject lines for its summer promotion: "Summer sale: 20% off all garden furniture" against "Is your garden ready for summer?" The platform sends each to 3,000 random addresses at 10 o'clock. After six hours version A counts 190 unique clickers and version B 140. That difference is large enough to act on, and the webshop sends version A to the remaining 24,000. For the next campaign the team notes the hypothesis "a concrete benefit in the subject beats a question", and tests it again with a different product to see whether the rule holds.

Frequently asked questions

Can I test more than two versions?

You can, but every extra version shrinks the groups and weakens the test. With a smaller list, stick to two.

What do I test after the subject line?

The sender name (company or person), the send time and the first button in the mail. One at a time, in the same way. On our page about sending newsletters you can read how to set that up in the campaign tool.

Can I test subject lines on transactional mail?

Better not. An order confirmation should be recognisable, not surprising. Save the testing for campaigns.

#a/b test subject line#test subject lines#email split test#a/b test sample size#email a/b testing
Call us
Send an email