I've looked at companies where two people can each pull a defensible retention number out of the same data and arrive at opposite conclusions. One version puts them in the top decile of anything. The other puts them on fire.

Both calculations were correct. Neither was true.

That gap is where a lot of companies live, and it isn't a spreadsheet problem. It's what happens when you take metrics built for one revenue model and apply them to a different one, then compare the output to benchmarks built for the first model.

The shape of the problem

Some businesses have revenue that moves on a cycle the customer controls, not one you control.

The periodic major service. The certification that comes due on a multi-year cycle. The comprehensive version that gets done once and then updated cheaply until it's time to do it properly again. Often there's a subscription underneath it, doing the work of holding the relationship between the big events.

Now run annual retention math across that. A customer who did the expensive thing last year and the cheap thing this year shows up as a 70 to 80% contraction. They didn't contract. They did exactly what they're supposed to do, on exactly the schedule the service is designed around. The number says something went wrong. Nothing went wrong.

Run it the other direction and it's just as bad. In a year where a lot of customers hit their expensive cycle, revenue expansion looks spectacular. Nobody sold anything. The calendar came due.

So you get panic and celebration on alternating years, both unearned, and the people looking at the dashboard slowly learn that the dashboard doesn't mean anything.

The second problem is quieter and worse

Cohort windows.

Most retention calculations compare a base period to a current period. Fine, as long as everyone in the cohort was present for the whole base period. They usually aren't. A customer who first showed up two months before the base window closed has two months of base revenue being compared against twelve months of current revenue.

That's not expansion. That's the calendar again, wearing a different costume.

In the cases I've looked at, the majority of measured expansion came from customers in exactly that position. Strip them out, hold the cohort to customers present for the full base year, and the number falls far enough to put every segment below 100%. And watch what happens to your best-performing segment: often it turns out to be one nearly-dead account that got re-sold into a much larger deal, single-handedly carrying the number.

The report wasn't lying. It was answering a different question than the one everyone thought it was answering.

Why benchmarks make it worse

A number you can't compare to anything gets interrogated. A number you can compare gets filed.

Tell a founder their gross retention is a little under the benchmark and they'll decide they're slightly behind and move on. Tell them their net revenue retention is well above it and they'll feel great and stop looking. In both cases the benchmark converted a question into a conclusion, and the conclusion was wrong because the number underneath it was measuring the wrong thing.

This is why "what's good?" is usually the wrong first question. The right one is "what is this number actually made of?"

What to do instead

Separate the revenue types before you measure anything. If you have a genuinely recurring product and a periodic service, they are two different measurement problems living in one QuickBooks export. Blending them produces a number that describes neither. Measure the recurring thing with recurring math. It's legitimate there, and whatever it says is real.

Measure the cyclical thing against its own cadence. Not annual retention. Of the customers who were due this period, how many came back? How long after due? At what level? That's the honest question, and it has nothing to do with a calendar year.

Turn the alarm condition into a queue, not a rate. If the failure mode is a customer not doing the thing when they're supposed to, then the metric is a count of who is past due and by how long. That's a list of names. Teams can work a list of names. Nobody has ever worked a percentage.

Check your cohort windows before you believe your expansion number. This one takes an afternoon and it's the highest-return afternoon on this list. Find out how much of your reported growth comes from customers who weren't there for the full base period.

Benchmark against yourself. Your own history is the only comparison that shares your business model. Industry benchmarks are useful for orientation and dangerous for conclusions, and the further your model sits from the one the benchmark assumes, the more dangerous they get.

The part that isn't about metrics

Underneath all of this is a simpler thing: knowing how your customers actually operate.

If your customer's buying behavior runs on a multi-year cycle, then your revenue runs on a multi-year cycle, and any measurement that assumes otherwise will generate false alarms and false victories forever. You cannot fix that with a better formula. You fix it by understanding the rhythm of the thing you sell and building the measurement around it.

The useful side effect is that once you know the rhythm, you know something much more valuable than a retention rate. You know when each customer is due. You know who's overdue. You know which customers are sitting in the long quiet stretch between the visible events, which is exactly when they start to wonder what they're paying for, and exactly when a subscription that was supposed to hold the relationship stops holding it.

That's not a dashboard. That's a work queue, a revenue forecast, and a churn early-warning system, and it usually comes out of data the company already has.

The test

Take your headline retention number and ask two questions.

What's in the numerator that isn't the same kind of revenue as what's in the denominator?

And how many customers in this cohort weren't around for the whole base period?

If either answer is uncomfortable, you don't have a retention problem yet. You have a measurement problem, and you can't tell whether you have a retention problem until you fix it.