The 40% benchmark is well established. The hundred responses it needs are the part nobody mentions — here is how to run it honestly at small scale.
The Sean Ellis test asks users how they would feel if they could no longer use your product. If 40% or more say "very disappointed", you have product-market fit. The threshold is well established. The sample size it requires is the part nobody mentions.
The follow-up questions are more useful than the score at any sample size.
One multiple-choice question, three options: very disappointed, somewhat disappointed, not disappointed. The 40% threshold emerged from Sean Ellis's work across a number of early-stage startups, where products above it tended to find sustainable growth and products below it tended to stall.
Two conditions matter and get dropped in most retellings. Survey only people who have used the core of your product recently and more than once — surveying signups who never activated measures your onboarding, not your fit. And ask for a reason immediately afterwards, because the reason is where the value is.
The score is a proportion, and proportions from small samples move violently.
| Responses | One extra "very disappointed" | Reading |
|---|---|---|
| 14 | Moves the score ~7 points | Noise with a decimal point |
| 40 | Moves the score 2.5 points | Directional at best |
| 100 | Moves the score 1 point | The benchmark holds |
At 14 responses, whether you clear 40% can come down to one person having a good week. Founders then act on that — pivoting something that worked, or scaling something that did not.
There is a second problem specific to small samples: the people who respond to your survey are your most engaged users, because they are the ones who open your emails. At a hundred responses that bias is diluted. At fourteen it dominates.
Three adjustments make the survey worth running before you have a hundred users.
Report a count, not a percentage. "Eleven of nineteen active users said very disappointed" is honest. "58% product-market fit" implies precision the sample cannot support.
Weight the follow-up question. Ask what they would use instead, and what specifically they would miss. Those two answers do most of the work — the small-scale PMF framework covers reading them.
Segment even at small numbers. If seven of your nineteen are freelancers and five of those seven said very disappointed, you have found a segment, not a score. That finding is worth more than clearing 40% overall.
Four questions. The first is the famous one; the rest are where the useful material is.
Question four is the most underused. Your users describe your ideal customer more accurately than you do, and their answers feed straight into positioning. Ask them in the wording covered by the Mom Test so the answers stay factual.
Send it to people who used the product at least twice in the last two weeks. Surveying your whole list produces a lower score that measures activation rather than fit, and you will draw the wrong conclusion from it.
Above 40% with a hundred responses, the standard advice applies: stop changing the product and start scaling acquisition. That is genuinely the moment to shift effort to building a repeatable channel.
Below 40%, the useful move is almost never a rebuild. Look at which segment scored highest and narrow toward it. A 30% overall score that is 55% among one customer type is a positioning instruction rather than a failure.
And if most respondents say they would use a specific competitor instead, your fit is fragile regardless of the percentage — that answer is telling you the product is substitutable.
A one-question survey asking users how they would feel if they could no longer use your product, with three options: very disappointed, somewhat disappointed, not disappointed. Products where 40% or more choose 'very disappointed' have historically found sustainable growth.
Around 100 from users who have used the core product recently and repeatedly. At 14 responses a single answer moves the score by about seven points, so whether you clear 40% can come down to one person having a good week.
Yes, but report a count rather than a percentage and weight the free-text answers. 'Eleven of nineteen active users said very disappointed' is honest; a percentage implies precision the sample cannot support.
People who used the product at least twice in the last two weeks. Surveying your whole list including inactive signups measures onboarding rather than fit, and produces a lower score that leads to the wrong conclusion.
Rarely a rebuild. Look at which segment scored highest and narrow toward it — an overall score of 30% that is 55% among one customer type is a positioning instruction. If most respondents name a specific competitor as their alternative, your fit is fragile regardless of the number.
Bring your responses, however few. Marcus tells you what the segment split says and whether to scale or narrow.
Try GhostCoach free →14-day free trial · cancel anytime · 30-day money-back on Lifetime