Field Notes

Why NDR Is the Most Misunderstood Metric in SaaS

A bit of history on why this number matters so much now

Board reporting didn't always look like this. Before about 2015, the metric that got tracked was logo churn: if you lost 5% of your customers, you reported 5% churn and moved on. That approach never told anyone whether the customers who stayed were actually growing their spend or shrinking it, which turned out to matter a lot. When COVID hit and new-logo acquisition basically froze for a while, boards needed a different answer to "how is this business actually going to grow now that we can't just keep adding new customers." Net Dollar Retention, how much of last year's revenue you still have, plus or minus whatever expanded or contracted, became that answer. It's been the dominant board-level metric ever since, through a few different valuation cycles, and it doesn't look like that's changing anytime soon.

Why a single number ends up getting misread

Part of why NDR gets so much attention is that it's one clean number, easy to put on a slide. That's also exactly why it's so easy to misread. It's an aggregate of churn, contraction, expansion, product adoption, customer health signals, NPS, and a handful of other things: pricing changes, new features, how long it takes a customer to get value, support capacity. One number can't tell you which of those is actually responsible for the result. A board sees NDR at 95% and assumes there's a churn problem, when the real issue might be that expansion has stalled out, or that one big account is masking what's happening to ten smaller ones. This is basically why I think about retention in terms of the five drivers I lay out on the Frameworks page. NDR is what comes out the other end. It's not the thing you're actually fixing.

The blind spot that catches a lot of people

Here's the one I see missed most often. NDR can look genuinely strong, 120% even, while a meaningful chunk of the customer base is quietly churning, as long as a few large accounts are expanding fast enough to cover for it. That's not the same thing as broad product-market fit across your whole customer base. It's more like enterprise fit specifically, showing up in the aggregate number as if it applied everywhere.

Where GDR comes in

This is where Gross Dollar Retention is useful, and it's worth explaining what it actually is, because it gets confused with NDR a lot. NDR includes expansion right alongside retention, so a company can lose real customers and still post a healthy NDR number if the customers who stayed are buying more (more seats, higher tiers, add-ons). GDR strips expansion out of the calculation entirely. It only measures how much of last year's revenue you kept, counting contraction and churn against you, with no credit given for upsells to offset the damage. A company can't paper over a bad GDR with a good upsell quarter the way it can with NDR. That's why I think of GDR as the more honest of the two numbers: NDR tells you whether the business is growing overall, GDR tells you whether you're actually keeping the customers you already have. Looking at both side by side matters too. If there's a wide gap between a strong NDR and a weaker GDR, that gap itself is usually telling you that a small number of accounts are doing most of the work.

The other way this number can mislead you

There's a second version of this problem worth knowing about, and it actually has a name: Simpson's Paradox, a pattern in statistics where a trend holds true in every individual group but reverses or disappears once you combine the groups together. Tomasz Tunguz has demonstrated this directly with NDR: using the same cohort dataset, he shows that averaging each cohort individually, totaling the whole year, and averaging by quarter produce three different NDR figures from identical underlying numbers. That same mechanism is what makes hypergrowth companies especially vulnerable to a subtler version of the problem: newer cohorts in their early expansion period can be large enough in dollar terms to mask slow, steady decay happening in older, more mature cohorts. So the blended, aggregate NDR number goes up, even while the company's longest-tenured customers are quietly leaving. The fix isn't a different metric. It's looking at the data differently, using a cohort retention table read left to right, to check whether older cohorts are flattening out into something sustainable or steadily decaying underneath a number that looks fine on the surface.

What this actually looks like in practice

At Sharpen, NDR was sitting at 70% when I joined, meaning the company was losing nearly a third of its revenue base every year before any new sales even came into the picture. The instinct in a situation like that is usually to look at churn first, since a low NDR feels like a churn problem on the surface. But going driver by driver, the bigger issue was execution: renewal and expansion ownership was still sitting with Sales, with no dedicated function actually watching for risk signals or working accounts proactively. Once that ownership moved into a real CSM function, NDR went from 70% to 100% and GDR went from 60% to 96% over about nine months. The gap between those two numbers closing that much told us the gains weren't just a few accounts expanding while others quietly left. The underlying retention itself was actually getting healthier.

Hyperscience is a different version of the same idea. NDR sat at 110% sustained through a period where the company scaled its customer base from 16 to 120 accounts and its CS team from 6 to 70 people across 9 countries. Hypergrowth is exactly the environment where Simpson's Paradox shows up: a flood of new cohorts in their early expansion period can make a blended number look great even if something's quietly wrong underneath. In that case it held up because product adoption and time-to-value were being actively managed alongside the scaling, not left to chance while the headcount numbers did the talking.

Two different situations, two different drivers responsible. That's really the point. The number alone wouldn't have told you that.

Wrapping up

NDR isn't a bad metric to track. It's just one that gets treated as more self-explanatory than it actually is. Looked at by itself, it invites the kind of misreadings above. Looked at as the output of several separate, knowable drivers, and checked against GDR and cohort-level data rather than read in isolation, it turns into one of the more useful diagnostic tools available in SaaS.

If you want the longer version of how NDR became the metric boards obsess over, I wrote a companion piece on that history.

← Back to Insights