Willingness to pay: what counts as proof
A hundred people saying they would buy is still zero people who did. Five rungs of evidence, from “sounds useful” to “paid again” — and the rule that keeps the weak ones from adding up to a strong one.
Short answer
What is willingness to pay, and how do you prove it?
The most a customer would accept before walking away. Surveys estimate the number; only behaviour proves anyone will pay it. The evidence runs in rungs — said it was useful, named a price, signed something, put money at risk, paid — and the highest rung you have reached is what you know.
The short version
- Stated methods (Van Westendorp, Gabor-Granger, conjoint) all sit on one rung: they measure what people say.
- Revealed evidence — deposits, pre-orders, a real checkout — is a different kind of observation, not a larger one.
- Weak evidence does not accumulate into strong evidence: a hundred survey yeses share one bias, all of it flattering.
- One paying customer proves the product is buyable. Three unrelated ones is the smallest thing that behaves like a pattern.
Free · No signup · Nothing leaves your browser
Count what people did, not what they said
For each kind of evidence, enter how many people have actually done it. Not how many might, not how many are in the pipeline — how many, so far.
- Rung 4 · Money received
Willingness to pay, demonstrated. Nothing else on this list does.
people - Rung 3 · Money at risk
Real money, refundable. The first rung that costs the buyer something.
people - Rung 2 · Commitment without money
Social commitment. Notably easy to sign and to forget.
people - Rung 2 · Commitment without money
That someone spent their own effort. Effort is cheaper than money.
people - Rung 1 · Stated intent
How the price is perceived. Saying yes in a survey costs nothing.
people - Rung 0 · Interest
That the topic is not offensive. Nothing about money.
people
Enter a number on any rung. If every rung is zero, that is an answer too — and a more useful one than most founders will admit to having.
Stuck on rung 1 because you have no price to test? Van Westendorp gives a believable range and Gabor-Granger ranks the prices inside it. Both stay on rung 1 — they are how you choose the number to put in front of someone, not evidence that anyone will pay it.
Said versus did
Every method for measuring willingness to pay falls into one of two families, and the difference between them is larger than the difference between any two methods inside a family.
Stated preference asks people what they would do. It is fast, cheap, and runs before the product exists — which is exactly why founders reach for it. It is also uniformly optimistic, because agreeing with a stranger about a hypothetical costs nothing.
Revealed preference watches what people actually do when something is at stake. It is slower and requires something real to put in front of them. It is also the only family that answers the question.
| Method | What it gives you | Rung |
|---|---|---|
| Van WestendorpYou have no idea what the number should be. | A range of prices people find believable | Rung 1 — stated |
| Gabor-GrangerYou have candidate prices and need to rank them. | Demand and revenue curves across prices you choose | Rung 1 — stated |
| Conjoint analysisYou have a research budget and a real design. Heavy, and still stated. | Trade-offs between features and price, with competitors present | Rung 1 — stated |
| A price on a real pageAlways, as soon as you can. This is the only one that leaves rung 1. | Who clicks buy, and who does not | Rung 3-4 — revealed |
Why weak evidence never adds up
The instinct is to score: a conversation is worth one point, a letter of intent five, a deposit twenty, and enough conversations eventually equal a deposit. Every prioritisation tool works this way, and here it is wrong.
The reason is that the errors are not independent. A hundred people saying they would buy are not a hundred separate observations that average out to the truth. They share one bias, pointing one way: agreeing is free, and people are kinder to a stranger with an idea than they are to their own wallet. Averaging more of the same systematic error does not cancel it — it only makes the number look substantial.
A deposit is not a bigger version of a survey answer. It is a different kind of observation, taken under different incentives, and that is the entire reason it counts.
Hence the rule the tool above applies and states out loud: your verdict is the highest rung you have reached. The number of people on every rung below it changes nothing — and is shown anyway, because it is usually what made the idea feel validated.
Four ways founders fool themselves
❌ Counting enthusiasm as demand
Forty people said the idea was great. None of them were asked for money, so the forty tells you the idea is not offensive — which is true of almost every idea that fails.
✓ Instead: Ask the next enthusiastic person for a deposit. The reply takes a day and settles more than another forty conversations.
❌ Treating a letter of intent as a sale
An LOI costs the signer nothing and commits them to nothing. It measures politeness under low stakes, which correlates with revenue far less than founders hope.
✓ Instead: Ask what would have to be true to turn it into a paid pilot, and put a date on it. The answer is usually more informative than the letter.
❌ Surveying your way to certainty
Stated-preference methods are useful and cheap, and they all live on the same rung. Running three of them gives you three views of what people say, not one piece of evidence about what they do.
✓ Instead: Use one survey to choose the number, then spend the rest of the effort putting that number in front of someone.
❌ Reading one paying customer as a market
One buyer proves the thing is buyable. It does not prove anyone else will, and the second sale is where most products discover they were a favour.
✓ Instead: Get to three, from people who do not know you. Then the pattern is worth extrapolating.
One of three critical criteria
In our Go/No-Go rubric, will they pay sits alongside is the problem real and does a market exist as one of three criteria marked critical — the ones no amount of strength elsewhere can compensate for. A perfect team with perfect unit economics and nobody paying is a NO-GO, and the rubric says so.
A 15-minute voice session works through all seven criteria and ends with a written GO / WAIT / NO-GO and the reasoning behind it. If your ladder above stops at rung 1, that is the criterion it will spend the most time on.
15 min · free tier, no card
Frequently asked questions
What is willingness to pay?+
How do you measure willingness to pay?+
Why does weak evidence not add up to strong evidence?+
How many paying customers prove a market?+
Is a signed letter of intent good evidence?+
Is this tool free, and does my data leave the browser?+
Related guides