How to A/Z test cold email variants
How-to guide Updated 12 July 2026
How do you A/Z test cold email variants?
Add your variants to a campaign step in HotHawk, up to 26 per step, A to Z. The list splits evenly across them and every reply is logged per variant. Then ask Claude which one is winning on reply rate, and tell it to switch the flat ones off so your best line gets more of the list.
No credit card required.
Before you start
You need a HotHawk account and the HotHawk MCP server connected to Claude. It takes a couple of minutes and no code. Full setup is on the MCP server page.
Step by step
-
1
You ask Claude
Add three subject line variants to step one of my SaaS Q3 campaign.
Claude callsmcp__hothawk__campaign_steps_variants_createstepId SaaS Q3, step 1 subject 3 new subjects body same bodyYou now have four variants on that step, A to D. HotHawk splits the list evenly across them and starts logging replies to each one.
-
2
You ask Claude
A week in, which variant is pulling the most replies?
Claude callsmcp__hothawk__workspaces_analytics_listworkspaceId current workspace breakdown by variantClaude reads the numbers back per variant: reply rate and positive reply rate for A, B, C and D, side by side, with a clear read on which one is winning.
-
3
You ask Claude
Variant C is flat. Turn it off and send the rest to the winner.
Claude callsmcp__hothawk__campaign_steps_variants_updatevariantId variant C status offC stops sending. The remaining leads route to the variants still in play, so your best line gets more of the list instead of your weakest.
Guessing at copy is expensive
Everyone has an opinion on which subject line is better. The problem is that opinions do not book meetings, and the only way to know for real is to put both in front of the list and count the replies. That is a test, and most people skip it because setting one up feels like a chore.
HotHawk lets you run up to 26 variants on a single step, A right through to Z. That is not a gimmick. It means you can test four subject lines, or three whole openers, in one go rather than pairing them off over weeks. The list splits evenly across every variant, so each one gets a fair sample.
Let Claude read the scoreboard
The setup is only half of it. The other half is reading the numbers and doing something about them, which is where a test usually dies in a spreadsheet nobody opens. Run it from Claude and that part gets easy. Ask which variant is pulling replies and it reads the data back per variant, then you tell it to switch the weak ones off in the same breath.
Here is the part that matters for trusting the result. HotHawk scores variants on reply rate and positive reply rate, the answers that actually mean something. Open tracking is off by design, so you are never fooled by a variant with a great open rate and a dead pipeline. The number you optimise on is the number that pays.
Fold the winner back in
A test is only useful if it changes what you send next. Once a variant has clearly won, you have a line worth keeping. Turn the losers off and let the winner take more of the current list, then carry that copy into your next campaign so you are not starting from a blank page again.
The whole time, the sending machinery underneath does not change. Emails still go out across your warmed, rotated mailboxes, inside the window you set, whether Claude is running the test or you are. You are learning what to say. HotHawk keeps handling how it gets sent.
The HotHawk tools behind this
Every action here is a real endpoint. Claude reaches it over the MCP server, and you can hit the same endpoints directly from the REST API if you would rather build it into your own stack.
-
mcp__hothawk__campaign_steps_variants_createAdd a variant to a step, up to 26 per step -
mcp__hothawk__workspaces_analytics_listRead reply and positive reply rate per variant -
mcp__hothawk__campaign_steps_variants_updateToggle a weak variant off
Frequently asked
What is A/Z testing and how is it different from A/B?
A/B testing pits two versions against each other. A/Z means you can run up to 26 on the same step, A through Z. In HotHawk you add as many variants as you want to test, and the list splits across all of them at once, so you learn faster than running one pair at a time.
How does Claude decide which variant is winning?
It reads the real numbers, not a guess. HotHawk logs reply rate and positive reply rate per variant, and Claude reads those back per variant so you can see which line is actually earning answers. You can tell it what "winning" means for you, replies or booked calls.
You do not track opens, so what am I actually testing on?
Replies, which is the number that pays. Open tracking is off by design, because open rates are noisy and adding a tracking pixel can hurt your sending. So an A/Z test in HotHawk is scored on real answers from real people, not a metric that inflates itself.
Do I need to write code to run an A/Z test?
No. You add each variant and read the results in plain language over the MCP server. The same variant and analytics endpoints are in the REST API if you would rather build the test loop into your own stack.
Can Claude change my copy or turn variants off without me?
Only if you ask it to. You can have it flag the loser and wait, or tell it to switch a flat variant off outright. Either way, warmup and mailbox rotation stay on the whole time, so a test never runs at the cost of your sending.