How to A/Z test cold email variants

How-to guide Updated 12 July 2026

How do you A/Z test cold email variants?

Add your variants to a campaign step in HotHawk, up to 26 per step, A to Z. The list splits evenly across them and every reply is logged per variant. Then ask Claude which one is winning on reply rate, and tell it to switch the flat ones off so your best line gets more of the list.

No credit card required.

Before you start

You need a HotHawk account and the HotHawk MCP server connected to Claude. It takes a couple of minutes and no code. Full setup is on the MCP server page.

Step by step

  1. 1
    You ask Claude

    Add three subject line variants to step one of my SaaS Q3 campaign.

    Claude calls mcp__hothawk__campaign_steps_variants_create
    stepId SaaS Q3, step 1 subject 3 new subjects body same body

    You now have four variants on that step, A to D. HotHawk splits the list evenly across them and starts logging replies to each one.

  2. 2
    You ask Claude

    A week in, which variant is pulling the most replies?

    Claude calls mcp__hothawk__workspaces_analytics_list
    workspaceId current workspace breakdown by variant

    Claude reads the numbers back per variant: reply rate and positive reply rate for A, B, C and D, side by side, with a clear read on which one is winning.

  3. 3
    You ask Claude

    Variant C is flat. Turn it off and send the rest to the winner.

    Claude calls mcp__hothawk__campaign_steps_variants_update
    variantId variant C status off

    C stops sending. The remaining leads route to the variants still in play, so your best line gets more of the list instead of your weakest.

Guessing at copy is expensive

Everyone has an opinion on which subject line is better. The problem is that opinions do not book meetings, and the only way to know for real is to put both in front of the list and count the replies. That is a test, and most people skip it because setting one up feels like a chore.

HotHawk lets you run up to 26 variants on a single step, A right through to Z. That is not a gimmick. It means you can test four subject lines, or three whole openers, in one go rather than pairing them off over weeks. The list splits evenly across every variant, so each one gets a fair sample.

Let Claude read the scoreboard

The setup is only half of it. The other half is reading the numbers and doing something about them, which is where a test usually dies in a spreadsheet nobody opens. Run it from Claude and that part gets easy. Ask which variant is pulling replies and it reads the data back per variant, then you tell it to switch the weak ones off in the same breath.

Here is the part that matters for trusting the result. HotHawk scores variants on reply rate and positive reply rate, the answers that actually mean something. Open tracking is off by design, so you are never fooled by a variant with a great open rate and a dead pipeline. The number you optimise on is the number that pays.

Fold the winner back in

A test is only useful if it changes what you send next. Once a variant has clearly won, you have a line worth keeping. Turn the losers off and let the winner take more of the current list, then carry that copy into your next campaign so you are not starting from a blank page again.

The whole time, the sending machinery underneath does not change. Emails still go out across your warmed, rotated mailboxes, inside the window you set, whether Claude is running the test or you are. You are learning what to say. HotHawk keeps handling how it gets sent.

The HotHawk tools behind this

Every action here is a real endpoint. Claude reaches it over the MCP server, and you can hit the same endpoints directly from the REST API if you would rather build it into your own stack.

  • mcp__hothawk__campaign_steps_variants_create Add a variant to a step, up to 26 per step
  • mcp__hothawk__workspaces_analytics_list Read reply and positive reply rate per variant
  • mcp__hothawk__campaign_steps_variants_update Toggle a weak variant off

Frequently asked

What is A/Z testing and how is it different from A/B?

A/B testing pits two versions against each other. A/Z means you can run up to 26 on the same step, A through Z. In HotHawk you add as many variants as you want to test, and the list splits across all of them at once, so you learn faster than running one pair at a time.

How does Claude decide which variant is winning?

It reads the real numbers, not a guess. HotHawk logs reply rate and positive reply rate per variant, and Claude reads those back per variant so you can see which line is actually earning answers. You can tell it what "winning" means for you, replies or booked calls.

You do not track opens, so what am I actually testing on?

Replies, which is the number that pays. Open tracking is off by design, because open rates are noisy and adding a tracking pixel can hurt your sending. So an A/Z test in HotHawk is scored on real answers from real people, not a metric that inflates itself.

Do I need to write code to run an A/Z test?

No. You add each variant and read the results in plain language over the MCP server. The same variant and analytics endpoints are in the REST API if you would rather build the test loop into your own stack.

Can Claude change my copy or turn variants off without me?

Only if you ask it to. You can have it flag the loser and wait, or tell it to switch a flat variant off outright. Either way, warmup and mailbox rotation stay on the whole time, so a test never runs at the cost of your sending.

Send cold emails that get delivered. Never miss a positive reply.

Serious deliverability paired with the best reply management in the market.

Start your 7 day free trial

No credit card required.

Premium warmup

Join our premium warmup pool

We have over 50,000 Google and Microsoft mailboxes in the pool and we are opening to the public soon. Be first to know when it's open.

Special offer

Get 50% more sending, FREE.

Send 50% extra emails per month on any plan, every month for as long as you're with us. Enter your details and we'll email your promo code over.

Your new boosted limits

  • Starter 100,000 150,000
  • Scale 300,000 450,000
  • Infra 500,000 750,000

Applies to any plan. One per customer.