All guidesCampaigns

How to A/B Test Your Email (the Right Way)

Most brands A/B test wrong: too many variables, too small a sample, no real conclusion. Learn the disciplined way ZHS tests so wins actually compound.

OutcomeAfter this lesson, you can
  • Plan the campaign using the framework in this lesson
  • QA the offer, audience, creative, and measurement plan
PrerequisiteRead this firstHow to Use Customer Stories and UGC in Your Emails

Why do most email A/B tests tell you nothing?

Most tests fail because they change several variables, stop before enough meaningful outcomes occur, or optimize a proxy instead of the business result. The ZHS operating rule is to isolate one decision, define the decision metric before launch, and collect enough evidence to make a reversible or durable choice with confidence.

Single
decision per diagnostic test
Pre-set
decision rule before launch
Business
outcome metric, not a vanity proxy

Most tests fail because they break the same discipline: isolate the decision, declare what would change your mind, and avoid calling a winner from a thin or unstable result. A small gap alone is not evidence; its meaning depends on the metric, baseline, audience, and number of observed outcomes.

The short version

Change one decision at a time so you can explain the result. Set the decision rule and primary metric before launch, then use the platform's analysis and the practical cost of being wrong to decide whether the evidence is sufficient. Test high-leverage questions before cosmetic refinements, use reliable downstream behavior alongside opens, and document both wins and inconclusive results.

A/B testing replaces opinions with data. That is the whole point.

But most brands run tests that prove nothing: several variables change at once, the audience is too small for the decision at stake, or a temporary lead is called a winner.

That is not testing. That is guessing with extra steps.

Done right, testing compounds. Each real win becomes your new control, and the next test builds on top of it.

That is how a program gets better every month instead of drifting on gut feel.


How many variables should you test at once?

For a diagnostic test, isolate one decision. If the subject line and the hero image both change and version B wins, you cannot say which change caused the result. Keep one element as the control, then use the result to decide the next test.

This is the whole game.

Change the subject line and the hero image in the same test, and version B wins. Now you have no idea which change did it.

You cannot ship the win because you do not know what the win was.

Here is what this looks like in practice.

One change. One clean read. Then the winner becomes your control and you test the next thing against it.


What should you test first in an email?

Start with the decision most likely to change the reader's willingness to engage or buy: audience, offer, angle, or message framing. Test implementation details later, once the larger decision is understood. The ordering is a ZHS prioritization heuristic, not a universal impact ranking.

Test the things that move the biggest numbers first.

A subject-line test can change the first impression of a send; a CTA test changes the next action for people who reach it. Treat both as hypotheses and judge them with the relevant downstream behavior.

Start high in the funnel and work down.

What to testWhy it is often worth testing
Audience, offer, and angleThey can change relevance and purchase intent.
Subject line and preview textThey shape the first inbox impression.
Send time and dayThey can change when an eligible reader encounters the message.
Hero section and first imageThey shape the next step after the open.
CTA copy and placementThey refine the action for engaged readers.

Work from the biggest unresolved business decision toward the smaller refinements. A button-color test is rarely the first priority when the audience, offer, or message promise is still unsettled.

The rule: match the metric to the variable, or you will crown a false winner.

Pick the metric that matches the variable

Choose a primary metric that is causally close to the decision, then check guardrails. For example, an offer or CTA test can be judged on reliable click, conversion, and revenue behavior; opens are a secondary inbox signal rather than the only verdict.


What is a valid A/B test sample size?

A test is valid when its analysis plan is set before launch and it has enough observed outcomes to distinguish a decision-worthy change from normal variation. There is no universal send count, conversion count, or confidence cutoff that fits every email program; required evidence depends on the baseline, metric, audience, and cost of a wrong decision.

A gap between two variants means nothing until you have volume behind it.

Use your platform's testing method or a documented analysis method, but do not let a dashboard label replace judgment. Before launch, write down the smallest change that would be worth operationally adopting, the metric you will read, the guardrails you will protect, and the date or traffic threshold at which you will decide.

If the test remains inconclusive, treat that as useful information: keep the current control, preserve the result in the log, and prioritize a higher-leverage question.


Should you test flows or campaigns differently?

Yes. Campaigns are one-shot: you split a single send, read the result, and the test is over. Flows are always-on, so version A runs against version B inside the automation and the sample builds itself over weeks. Because a flow runs forever, even a small win keeps paying out every day.

Campaigns are one-shot. You split the send, read the result, and the test is over. Good for subject lines, offers, and send times where you have list volume in a single blast.

Flows are always-on, so they test differently and better. You set version A against version B inside the automation, let real traffic split over weeks, and the sample builds itself.

Because a flow runs forever, even a small win keeps paying out every day.

Campaign tests
  • One send, one read
  • Best for subject line, offer, send time
  • Needs list volume in a single blast
Flow tests
  • Runs continuously, sample builds over time
  • Best for hero, CTA, email order, timing
  • Small wins compound every day

Why should you document A/B test results?

Because a test you do not write down is a test you will run again in 6 months. Keep a simple log: what you tested, the two versions, the numbers, the confidence level, and what you shipped. That log stops you re-testing settled questions and turns scattered wins into a playbook new team members can read.

A test you do not write down is a test you will run again in six months.

Keep a simple log: what you tested, the two versions, the numbers, the confidence level, and what you shipped.

That log does two things. It stops you from re-testing settled questions, and it turns scattered wins into a playbook.

When a new team member asks why your emails send at 8am, the answer is a line in the doc, not a shrug.

Process

Why should you document A/B test results

  1. 01
    First testSubject line

    Isolate the decision, define the primary metric and guardrails, then log the outcome even if it is inconclusive.

  2. 02
    Ship and logDecision becomes the control

    Adopt only the outcome you can explain, and record the audience, dates, metric, and rationale.

  3. 03
    Next testOffer or angle

    Run the next question against the documented control so the learning remains cumulative.

  4. 04
    ContinueSend time, hero, CTA

    Retest only when the audience, season, offer, or evidence changes enough to reopen the question.

Structured from the canonical article steps for responsive reading and presenter mode.


When should you not run an A/B test?

Do not run a formal split when the available audience cannot reasonably answer the decision, when the change is not safely reversible, or when a test would disrupt an important moment. Use a documented production standard, qualitative research, or a longer-running flow test instead, then revisit the question when the evidence base improves.

Low volume kills tests

If a split cannot produce decision-quality evidence in the relevant window, do not force a winner. Preserve the question and use the safest reasonable control.

Small lists should not A/B test every send. You will never reach significance, so you end up making changes on luck.

Instead, apply proven best practice, grow the list, and save formal testing for the moments where you have the volume to get a clean read.


What are the most common A/B testing mistakes?

The recurring failures are changing several things at once, choosing the metric after seeing the result, treating an unstable lead as a winner, testing without enough evidence, failing to document the outcome, and polishing trivia before resolving the audience or offer. Fix those and the program can learn without overclaiming.

  1. Changing two things at once. You cannot tell which one won, so you cannot ship the win.
  2. Calling it too early. Use the pre-declared decision rule, not a first-hour lead.
  3. Measuring the wrong metric. For a matched sales-email split test, judge an offer or CTA with the reliable behavior it is meant to change, such as revenue per recipient, and keep meaningful guardrails visible.
  4. Testing on tiny lists. No volume means no significance means no real answer.
  5. Not documenting results. Undocumented wins get re-tested and quietly lost.
  6. Testing trivia first. Optimize opens and offers before you touch button color.

Get Expert Help

Our team runs disciplined tests across dozens of DTC brands, so we know which variables move revenue and which just waste a send.

If you want a program built on real data instead of guesswork, we can help.

See our pricing | Apply to work with us

Need help implementing this?

We build and manage complete email & SMS programs for DTC brands. Get a custom plan for your brand.

Apply Now

Join 2,000+ ecommerce strategists

Get all my brand breakdowns, Klaviyo guides, and the exact systems behind $50 million in DTC sales, directly in your inbox.

We respect your privacy. Unsubscribe anytime.