Product-Market FitSaaSB2BMeasurement

Product Market Fit for SaaS: A Practical Guide to Measuring It

Product-market fit for a SaaS product is repeatable, durable alignment between your product, a segment you can name, and a job that segment actually needs done. Here is how to measure it with the disappointment survey and the metrics you already have, without inventing benchmarks that do not exist.

AR
Anton Reed
· 18 min read

For a SaaS product, product-market fit means a segment you can describe keeps choosing your product, reaching its value, paying for it, and coming back, because the job it does for them is material. Not because switching is annoying.

That is a judgment built from several kinds of evidence, and no single number issues the verdict. The most useful early instrument is the disappointment survey: ask users how they would feel if they could no longer use your product, and count the share who say "very disappointed." 40% is the benchmark to aim for, and later in this article you will see exactly how firm that line is and is not.

This guide covers what fit means for a SaaS business, what the survey does and does not tell you, how to read your existing metrics without borrowing thresholds nobody validated, and what to do with the answers at different stages.

One warning about the genre first. Most articles on this topic hand you a table of universal benchmarks: churn under X, net revenue retention over Y, time to value under Z. Those numbers are almost always presented as settled facts about SaaS when they are not. Self-serve products at $19 a month and enterprise products on annual contracts do not share a churn scale. Where a number is worth having, this article names who produced it and what it rests on. Where no defensible number exists, it tells you what to compare against instead, which is usually your own trend.

What the disappointment survey measures

One question produces the score:

"How would you feel if you could no longer use [product]?"

The answer options are "very disappointed," "somewhat disappointed," "not disappointed," and an N/A option for people who have stopped using the product. Your score is the percentage of eligible respondents who completed that question and chose "very disappointed." State that denominator every time you report the number, because "percentage of responses" quietly changes meaning depending on who was invited and who answered.

Sean Ellis created this survey and the 40% reference line. His own account of where the line came from is worth reading in full, because it is more modest than the way it usually gets repeated:

"Admittedly this threshold is a bit arbitrary, but I defined it after comparing results across nearly 100 startups. Those that struggle for traction are always under 40%, while most that gain strong traction exceed 40%."

Read the asymmetry in that sentence. The companies that struggled were always below 40%. Only most of the ones with strong traction were above it. Those are two different strengths of claim, and collapsing them into a single dividing line is where the number gets misused. A low score is a strong signal that something is wrong. A high score is encouraging evidence, not a certificate. 39% is not a verdict, and 41% is not an achievement unlocked.

In his later writing Ellis put it more softly still: it "becomes possible to sustainably grow a product when it reaches around 40% of users who try it that would be 'very disappointed' if they could no longer use it."

Why bother with an attitudinal question at all when you have usage data? Because it is a leading indicator. Ellis's stated reason for building it was that the alternative is to rely on instinct or to wait for retention cohorts to mature, which can cost months. The survey buys you an earlier read. It also buys you the free-text answers, which are the part that actually tells you what to build.

What it does not measure

The survey measures attitudinal necessity: whether the people who have genuinely used your product would feel a loss without it. That is one layer of evidence. It says nothing directly about whether they will pay, whether they will still be here in six months, whether you can find more of them repeatably, or whether the market is large enough to matter. Ellis himself flagged the last one in the original post: progressing beyond early traction "requires that these users represent a large enough target market to build an interesting business."

So treat a good score as one converging line of evidence. Fit needs several:

  • Market truth. A material job, real urgency, budget and authority, known alternatives, a reachable market.
  • Attitudinal necessity. The disappointment score, plus why, plus the main benefit people name.
  • Behavioral value. Activation, completion of the core job, repeat use at the job's natural cadence, cohort retention.
  • Commercial value. Payment, conversion, renewal, expansion, and where your economics are heading.
  • Pull and repeatability. Organic inbound, referrals, an ICP and an onboarding path you can run again without founder rescue.
  • Delivery durability. Whether you can actually serve these customers at volume, and whether the position holds as the market moves.

Acquisition shows that a promise can attract attention. Retention and repeat purchase show whether the promise was kept. Both matter, and neither substitutes for the other.

How many responses you need

There is no universal count that makes a score valid, and anyone who tells you otherwise is repeating a convention as if it were a rule. Two primary figures are worth knowing, and they belong to two different people:

  • Ellis: "a minimum of 30 responses is needed before the survey becomes directionally useful. At 100+ responses I am much more confident in the results."
  • Vohra, from Superhuman's write-up: "you start to get directionally correct results around 40 respondents, which is much less than most people think."

Cite whichever you are using. Do not average them into a number neither of them said.

Then report the count next to the percentage, because a percentage on its own hides how much you actually know. Illustrative arithmetic: 16 "very disappointed" answers out of 40 respondents is a 40% score, and an idealized 95% interval around it runs from roughly 26% to 55%. That interval assumes your respondents are a random sample of the population you care about, which they are not. It does nothing about the fact that engaged users answer surveys and quiet ones do not. So the real uncertainty is wider than the arithmetic suggests, and the honest reading of "40% on 40 responses" is "worth acting on, not worth announcing."

Smaller samples still earn their keep. Ellis notes the survey "can still be insightful with only a few responses if at least some of them would be 'very disappointed'," because the free text from those people tells you what your product is actually for. Just do not build a trend line out of eight answers.

Reading your SaaS metrics without borrowing thresholds

Your product analytics and billing data are the behavioral and commercial half of the evidence stack. The failure mode is not measuring them, it is grading them against numbers someone published for a different kind of company.

Retention and churn

Churn is a behavioral signal, and an important one. It is not a direct measure of fit, and there is no universal monthly rate that certifies or condemns a product. The number moves with price point, contract length, buyer type, self-serve versus sales-led motion, and how much of your base is on a trial-shaped plan. A rate that would be alarming for annual enterprise contracts can be ordinary for a low-priced self-serve tool bought on impulse.

Four habits make churn readable:

  1. Compare cohorts against themselves over time. Does the January cohort's month-three retention look better than the October cohort's month-three retention? That comparison is yours, it controls for your own market and price point, and it answers the question you actually care about: is the product getting better at keeping people?
  2. Separate logo churn from revenue churn. Losing many small accounts and losing a few large ones produce very different revenue pictures and call for different responses. Reporting only one hides half the story.
  3. Segment it. Losing free trial users who never activated tells you something about qualification and onboarding. Losing paying power users who activated months ago tells you something about the product. They are not the same event, and averaging them produces a number that describes nobody.
  4. Treat a sharp change as the signal, not the absolute level. A retention curve that flattens is the encouraging shape. A curve that bends downward after a pricing change, an ICP shift, or a release is the thing to investigate, whatever the headline rate happens to be.

Attitudinal decline and behavioral decline are related, and the attitudinal one often shows up first, which is the reason to run the survey at all. Anyone who tells you the behavioral consequence arrives on a specific schedule is guessing.

Net revenue retention

Net revenue retention is expansion minus contraction minus churn, measured against your starting revenue from existing customers. Formula: (starting MRR + expansion - contraction - churn) / starting MRR x 100.

NRR below 100% means the revenue you keep and grow from existing customers is not covering what you lose from them. That is a real warning sign and worth understanding rather than explaining away.

It is not proof that you lack product-market fit. Two reasons. First, early companies with genuine fit in a narrow segment routinely run below 100%, because they have few accounts, no expansion motion yet, and one departure moves the number several points. Second, NRR reflects pricing and packaging as much as product value. A product people would hate to lose can still show weak NRR if there is nothing to upgrade to, if seats are the only expansion lever and teams are not growing, or if the plan structure caps how much a delighted customer can spend. Fixing packaging can move NRR without changing the product at all, which should tell you how much of the metric is about product love.

Read NRR the same way as churn: as your own series over time, decomposed into its parts. Rising expansion revenue from accounts that started small is a genuinely strong signal, because those customers voted twice.

Unit economics

Acquisition cost against lifetime value tells you whether the business can afford the market, not whether the market wants the product.

Illustrative arithmetic, hypothetical numbers: if it costs $100 to acquire a customer who pays $29 a month and stays 14 months, that is $406 of gross revenue per customer against $100 of acquisition cost, a ratio of about 4:1. The commonly repeated target is 3:1 or better. Treat that as an industry convention rather than a validated finding, and notice how much of it depends on inputs you are estimating: average lifetime is a forecast, especially before you have cohorts old enough to have finished churning, and using revenue rather than gross margin flatters the ratio.

The trend in your own ratio is more informative than its level. CAC rising while lifetime holds flat usually means you have exhausted the easiest part of your market, which is a positioning problem before it is a spend problem.

Activation and time to value

Getting users to the core value quickly matters. There is no universal target for how quickly, and "it should happen in minutes" is a claim about a category of product, not about SaaS generally. A tool that reformats a spreadsheet should deliver value in seconds. A product that has to ingest a quarter of billing data before it can say anything useful cannot, and that is not a fit problem.

Define your activation event as the moment a user completes the core job at least once, then track two things against your own history: the share of new signups who get there, and how long it takes them. Improving those numbers is nearly always worth doing. Grading them against a benchmark from someone else's product is not.

Growth rate

Growth is evidence about your acquisition engine. Whether it is evidence about fit depends entirely on where it comes from and whether it survives. Paid acquisition that fills the top of a funnel which empties again in three months is not fit, and no growth rate band will tell you which situation you are in. Look instead at what fraction of new revenue arrives without paid spend, whether cohorts are retaining better over time, and whether customers are sending you other customers.

Running the survey

1. Decide who is eligible

Only survey people who have had a fair chance to experience the core value. That is a behavioral definition, not a time-based one, and it depends on your product.

Both primary sources set the bar this way. Ellis: limit the sample to people who have had "real usage" recently, and for a ride-sharing product that means people who took a ride, not people who downloaded the app. Superhuman used "those who used the product at least twice in the last two weeks."

Write your rule down as an explicit filter: the core action that counts, how many times, within what window, on which product version, and which roles you are excluding. Then keep that rule stable, because changing eligibility changes the score independently of anything you shipped.

2. Ask the follow-ups too

The score tells you where you stand. The free text tells you what to do. FitSignal's PMF survey has seven fixed questions, and only the first is scored:

  1. How would you feel if you could no longer use [product]?
  2. Please help us understand why you selected this answer?
  3. What would you use if [product] were no longer available?
  4. What is the main benefit you receive from using [product]?
  5. What type of person do you think would benefit most from [product]?
  6. How can we improve [product] for you?
  7. What is your job title?

Question 4 is where your differentiation is described in customers' own words. Question 5 is where your positioning language comes from, because happy users almost always describe themselves. Question 3 gives you their perceived alternatives, which is useful and is not the same thing as your real competitive set. More on the design of each one: PMF survey questions.

3. Segment, but keep the raw number

An overall score can hide two different products. A hypothetical illustration: 62% among activated paying accounts and 14% among free trial users who never finished setup produces a mediocre blended number that describes neither group and points at no decision.

Useful cuts are plan type, activation depth, tenure, acquisition channel, company size, and job title. Report every segment's sample size alongside its score, and keep the raw baseline visible next to the segmented view permanently.

That last rule matters more than it sounds. Segmentation can raise a score by changing the denominator rather than by improving anything, and the same arithmetic will lift almost any score if you cut narrowly enough. A post-hoc lift is a hypothesis about where fit is stronger. Test it on a fresh eligible cohort before you believe it.

4. Decide your cadence deliberately

FitSignal does not prescribe a survey rhythm, and you should be suspicious of any article that does. The right interval is the one that matches how often people actually use your product. A tool used daily produces meaningful re-reads far sooner than one used at quarter close.

A first survey is a state, not a baseline. It is one reading of one cohort at one moment.

When you want a trend, make one choice explicitly and then keep it stable: are you measuring newly eligible users as they qualify, or re-asking the same people at an interval? Both are legitimate. They answer different questions, and mixing them silently makes the resulting line uninterpretable. Re-asking has a specific caveat worth stating plainly: answers from the same person are not independent readings, so movement in that series is harder to attribute to product change. In FitSignal, recurring measurement is an explicit setting for both PMF and NPS surveys, off by default, with the interval set in days.

Acting on the results

Find out who your product is genuinely for

Look at the very-disappointed group's answers to questions 4 and 5, and build a profile of the person your product already works for.

The framework worth using here is the High Expectation Customer, originated by brand strategist Julie Supan in First Round Review in 2016. Her definition: "The high-expectation customer, or HXC, is the most discerning person within your target demographic. It's someone who will acknowledge, and enjoy, your product or service for its greatest benefit."

Discerning, not demanding. That distinction does real work. A discerning customer recognizes what your product is best at and appreciates it, and is someone others in that market take seriously. A demanding customer generates the most requests. Those are different people, and building for the second one produces a longer feature list rather than a sharper product. Supan's argument for the narrow target: "If your product exceeds their expectations, it can meet everyone else's." More on identifying yours.

An HXC is also not simply everyone who answered "very disappointed." That group is where you look. The profile is what you build from it.

Split the roadmap two ways

Rahul Vohra's account of Superhuman's process gives the most usable prioritization rule in this literature. Half the roadmap goes to deepening what your very-disappointed users already love. The other half goes to removing what holds back the somewhat-disappointed users who named the same main benefit. His reasoning, verbatim: "If you only double down on what users love, your product-market fit score won't increase. If you only address what holds users back, your competition will likely overtake you."

Treat 50/50 as a starting heuristic rather than a theorem. The part that generalizes is the two-lane structure; the ratio is a dial you can justify moving for a quarter.

The filter on lane two is what keeps it from becoming a request queue. Somebody who is somewhat disappointed and wants your product to become a different product is not who you are building for. Somebody who named your main benefit and is blocked by one specific thing is. More on sequencing that work.

Read the score alongside behavior, without inventing the link

When fit strengthens, you would expect retention, expansion, and referrals to follow. Watch for that. What you cannot responsibly do is claim a fixed relationship between a survey score and a business metric, or a schedule on which one moves the other. Track them side by side, look for divergence, and treat divergence as a question rather than an error.

Divergence is often the most informative thing on the dashboard. A score that is rising while retention is flat can mean your newest cohorts differ from your older ones, that you changed eligibility, or that the people answering are not the people leaving.

The Superhuman example, accurately

This case study gets cited constantly and mangled almost as often. The sequence as Vohra reports it:

Superhuman's initial score in the summer of 2017 was 22%. He then segmented to the very-disappointed users who loved the product most, and reports that "our product-market fit score jumped by 10%. We weren't quite at that coveted 40% yet, but we were a lot closer with minimal effort." That is the move to 33%. Then the product work: "Within just three quarters of our work to improve the product, the score nearly doubled to 58%."

Two steps, two different mechanisms. The first is segmentation, which changed who counted. The second is three quarters of shipping against the two-lane roadmap. Collapsing them into one number, or describing the improvement as a result of tracking the metric, removes the only part that is actually instructive.

Vohra also makes a point most retellings drop: a score can fall as you grow. Early adopters forgive gaps that later users will not, so pushing past that group can reduce the number even while the product improves. If that happens, it is information about who you are now selling to.

What fit looks like at different stages

Stages are useful for deciding what evidence to go and get. They are not thresholds you clear.

Pre-launch or MVP. You are establishing that a specific job is material to a specific group, and that your approach is a plausible answer to it. Evidence: people describing the problem in their own words with urgency attached, existing workarounds you can name, and a handful of early users who complete the core job unaided. Survey a small group as soon as anyone has genuinely used the product, and read the free text rather than the percentage.

Early, one segment. You are looking for the pocket where the product already works. Evidence: a segment you can describe without hedging, users completing the core job repeatedly without hand-holding, and a very-disappointed group whose answers to "main benefit" agree with each other. That last agreement is the underrated signal. Scattered benefit answers usually mean you have several weakly-served use cases rather than one well-served one.

Growth, defending the core while expanding. You are testing whether the pocket generalizes. Evidence: newer cohorts retaining at least as well as older ones, activation improving as you widen the funnel, expansion revenue appearing without discounting, and segment-level survey scores that hold up as the mix changes. Compare each new segment against your core segment's own history, not against a published band.

Scale. You are managing fit across segments that are diverging. Evidence: retention and disappointment scores tracked separately per segment, a repeatable onboarding path that does not depend on you, and early warning when a segment's score or retention starts to bend. Expect fit to be uneven here, and expect at least one segment to be quietly getting worse.

Mistakes that are easy to make

Measuring people who have not used the product. A survey sent on signup day measures hope. Set an eligibility rule based on the core action, and enforce it.

Reading only the blended score. It can conceal a strong segment and a weak one at the same time. Segment, publish the sample sizes, keep the raw baseline.

Treating one reading as a baseline. It is a state. A trend needs a stable, declared method.

Averaging feedback across groups. The very-disappointed group tells you what to protect. The somewhat-disappointed group that shares their main benefit tells you what to fix. Those are different inputs to different decisions.

Writing off the not-disappointed group entirely. There is one question where deprioritizing them is defensible: what to build to deepen love in your core. For churn, pricing, positioning, support load, and market sizing, they are often the only people who can tell you what is wrong. They are the cheapest research you have.

Confusing growth with fit. Paid acquisition can produce a growth chart with no retained value underneath it. Ask what happens to a cohort after month three.

Borrowing benchmarks. This is the one that quietly ruins the rest. A number copied from an article about a different price point, buyer, and motion is not a target; it is noise with a decimal point. Your own trend, segmented and honestly denominated, is the comparison that means something.

Calling a survey score product-market fit. It measures attitudinal necessity in one declared cohort at one point in time. That is genuinely useful and it is one layer. Fit is the judgment you make when the layers agree. What else has to be true.

Where FitSignal fits

FitSignal runs the seven-question PMF survey, scores question 1 against the 40% line, and shows the trend, the distribution across the three answers, and a confidence level alongside the number rather than the percentage on its own. Word Cloud Analysis pairs two linked views: what the very-disappointed core says it loves, and the blockers named by somewhat-disappointed respondents tied to each of those benefits. The AI Improvement Analysis clusters free-text answers into themes ranked by frequency, severity, and persona weight, with each theme linked back to the verbatim quotes behind it. Personas and segments let you break the score down by group instead of shipping against a blended average.

Licensed NPS surveys are included in every plan, which is useful for a different question. The disappointment survey asks whether people would miss you; NPS asks whether they would recommend you. The two are not interchangeable.

The free plan covers one project and 250 collected responses a month through the in-app widget or a share link, which is enough to get a first read on a real cohort and decide whether the number is worth building a process around.

Run this exact survey on your product.
Free plan · 250 responses a month via widget or link · first score in days
Start free →