A note on the examples. In the particular is contained the universal. That is Joyce, and it is my defense for the fact that everything below happens to be about email. Email is where I have watched this most clearly and where the arithmetic is easiest to show.

The mechanism has nothing to do with email. It runs anywhere a team produces options that no outcome will act on, at a volume that could never have separated them: pricing experiments, onboarding sequences, landing pages, feature flags, pilot programs, agency creative. If your version is not email, substitute it. Nothing in the argument changes.

I have watched teams build four variants of an email, label them carefully, send them, and never once look at the results.

Not out of laziness. The variants got built, which is the expensive part. Somebody wrote them, somebody reviewed them, somebody set up the split. Then the campaign went out and nobody came back to read what happened, and the next campaign got built the same way.

That is not testing. It is the costume of testing, and it is worth being precise about what is missing.

A test needs a decision attached to it

The thing that makes a test real is not the split. It is not the labels, and it is not the language about letting the data decide.

It is that something you would otherwise do changes based on the outcome.

If no possible result would alter what you do next, you did not run a test. You sent four emails. The split added cost and produced nothing, because there was never a decision waiting on the other end of it.

Write down, before you send, what you will do differently if A wins and if B wins. If those two sentences are the same sentence, cancel the test and send your best one.

That question takes about a minute and it kills most of the tests people are running.

At low volume the readout is noise anyway

Here is the part that makes it worse rather than merely wasteful.

Split a few hundred sends across four arms, on a channel where a good outcome is a handful of responses, and the gap between the variant that got 2 and the variant that got 0 is a coin flip. It is not a small effect. It is not a weak signal. It is nothing, and no amount of staring at it will make it into something.

You would need thousands of sends per arm before those numbers separate from chance, and most campaigns will never have that.

Which means nobody reading the results was almost a mercy. A diligent reader would have found a difference, believed it, and changed the next campaign based on a number that was randomness with a label on it. The ritual was harmless only because it was ignored. Start taking it seriously without fixing the volume problem and it becomes actively dangerous.

The format destroys the signal you could have had

This is the part I had not thought through carefully enough until recently, and it is the real cost.

I have argued that convergence validates, not volume. If twenty people tell you the same thing, unprompted, in their own words, that is validated at twenty.

Now look at what a four-way split actually collects. People can click or not click. Open or not open. Reply or not reply. There is no mechanism in that data for anyone to say the same thing as anyone else, because nobody is saying anything. Scripted variants cannot converge. The format has no channel for agreement to appear in.

So at low volume you have chosen the one data type your sample is too small to read, and given up the one data type it is large enough to read.

Five replies is a useless sample for split-test arithmetic and a rich one for reading. What did they respond to. What did they push back on. What words did they use for the problem, which are almost never the words on your website. That is exactly the open-ended, in-their-own-words input that convergence can actually process, and it is sitting in the inbox while everyone is looking at a rate.

Sequential, not parallel

The model that fits small volume is not four things at once. It is one thing, then the next thing.

One best shot. Full send. Read every reply, including the annoyed ones, especially the annoyed ones. Then change the whole approach based on what came back and send again.

Most teams operating at small scale are already doing this informally. The campaign flops, somebody reads the silence, the angle gets rewritten, it goes out again. That loop is the actual methodology. The variants were decoration sitting on top of it, and the loop was doing all the work the whole time.

At that volume the channel is not a conversion machine. It is a customer development instrument that occasionally produces revenue, and it is a very good one, because strangers with no stake in your feelings will tell you things your customers have gotten too polite to say.

Either wire it or kill it

If you have accounts with genuine volume, wire a decision rule to the split before it goes out. Name the metric, name the threshold, name the change you will make. Then the test is real and worth its cost.

Everywhere else, stop producing variants. That work is not free. Someone is spending hours a week generating options that no result will ever act on, and those hours could go into reading replies, which is where the only readable signal at that scale actually lives.

What you cannot do is leave it as it is, because the current state has a cost and no benefit, and it is protected by how good it sounds. We tested it ends conversations. Nobody asks whether the test could have changed anything.

That is the same move as waiting for significance, arriving from the opposite direction. One postpones the decision until a threshold that never comes. The other performs the decision so thoroughly that nobody notices it was never made. Both are judgment deferred to a tribunal that does not convene, and both are more comfortable than saying the honest thing out loud.

Which is usually some version of: I think this one is stronger, I am not certain, and we are going to find out by sending it.