Can AI validate my startup idea?
We sell one of these tools, so start with the part that argues against us: a model reading your description is scoring your writing, not your market. Two founders with the same business get different verdicts depending on who writes better, and either of them can raise the score by editing the pitch instead of changing the company.
That objection is correct. It is also not the whole line. Below is where the line actually falls, what moves it, and what nothing moves.
Short answer
Can AI validate my startup idea?
Partly, and the part matters. An AI can check whether your reasoning holds together, who already sells this, whether your market arithmetic survives one question, and whether your customer questions are leading. It cannot establish demand, because demand is a fact about other people's behaviour.
Where the line falls
- An AI can settle four things from text alone: internal contradictions, who the competitors and substitutes are, whether the market arithmetic is defensible, and whether your planned questions are leading
- It cannot settle demand. The only instruments for demand are a payment, a deposit, a signed commitment, or an alternative someone cancelled
- A confident verdict on demand derived from a paragraph is scoring your writing: edit the pitch and the number moves while the business does not
- A 15-minute interview changes the input, not the laws — it captures what you did not think to write down, and exposes the gap between a claim and its evidence
- The useful output of any validator is not a score but an ordered list of what to go find out, cheapest question first
- No prompt can add a step that does not exist: being asked rather than typing, three models checking one another, a landing page real people sign up on, a survey that needs five real answers, and a record of how far the earlier prediction missed
- 42% of startups fail for lack of market need (CB Insights) — which is a reason to go ask people, not a reason to trust a score
What text can settle, and what it cannot
The split is not about how good the model is. It is about whether the answer exists in language at all, or only in someone else’s behaviour.
| Question | Who can answer it | Why |
|---|---|---|
| Whether your reasoning holds together | An AI can | Contradictions between your market size, your price and your channel are visible in the text itself. So is a plan whose first hundred customers arrive through "social media, SEO and word of mouth". |
| Who already sells this, including the boring substitutes | An AI can | Competitors are public. So are spreadsheets, agencies and doing nothing, which are the real incumbents and the ones founders forget to count. |
| Whether your market arithmetic is defensible | An AI can | A top-down number derived from a global market figure falls apart under one question about who is actually addressable. That question can be asked mechanically. |
| Whether your customer questions are leading | An AI can | A question that pitches, asks for an opinion, or asks about the future instead of the past is recognisable from its wording alone. That is a language problem, and language is what models are good at. |
| Whether anyone will pay | Nothing can, except payment | Demand is a fact about other people. The instruments are a payment, a deposit, a signed commitment, or an alternative someone cancelled. Enthusiasm about a hypothetical is not one of them. |
| Whether you will still care in a year | Nothing can | The most common way a small company dies is the founder losing interest, and no market analysis has ever predicted it. |
| Whether your channel works at your cost | Only running it | Channel economics are specific to your product, your price and your moment. A general estimate here is astrology with decimals. |
Why a score is the wrong output
A single number invites one question — is it high enough — and that question has no useful answer. Worse, it is editable: rewrite the description more confidently and the number rises while nothing about the business has changed. Any output you can improve without leaving your desk is measuring the desk.
The output that survives contact with reality is an ordered list of unknowns, where each item says what would count as an answer and roughly what finding out costs. Then validation stops being a verdict you receive and becomes a week you spend.
This is why every tool on this site shows how much to trust its own result: a range instead of a point, the share of an answer that was invented, the worst case next to the average. A tool that hides its uncertainty is not being confident, it is being quiet.
What a 15-minute interview changes
It changes the input, not the laws. A description contains what you thought to write down. Being asked follow-up questions for fifteen minutes contains what you did not, and a thin answer invites another question where a paragraph simply sits there.
Concretely, in our case: three consultants ask in turn, the conversation runs in 21 languages, competitors are pulled live rather than recalled, key claims are cross-checked by three independent models, and the verdict is computed from thresholds instead of composed by a model — so it cannot be talked into a kinder answer. The result is up to 20 reports, of which the summary is free.
And what it does not change: the conversation itself still does not measure demand. It tells you which of your assumptions is load-bearing and which question to go answer first. Measuring demand is the next step, and it is a step rather than a smarter analysis — which is the whole point of the section below.
What actually closes the loop
Most of the writing on this subject argues that a chatbot is too agreeable, and cites good research to prove it. The argument is true and it is also easy to answer: if the problem were the model’s manners, a better prompt would fix it, which is why prompt packs sell.
So here is the harder version. No prompt can add a step that does not exist. The difference between a conversation and a validation is how many of the steps touch the world outside the conversation.
The input to a chat is what you chose to write down. Fifteen minutes of adaptive follow-ups is what you did not, and a thin answer draws another question instead of passing.
Key claims are cross-checked by three independent models. One model agreeing with itself is not verification, however the prompt is worded.
The GO / NO-GO comes from thresholds rather than from text generation, which is why it cannot be talked into a kinder answer by a more confident description.
The brand consultant generates a landing page and it is published at its own address, where real visitors sign up and those signups are recorded. That is demand, not an opinion about demand.
Questions are generated from your own hypotheses and sent out as a public link. Nothing is scored until at least five real people have answered.
When those answers come back, the system computes how far its earlier prediction was from what people actually said, and stores that gap. A chat answer is never compared with anything, ever.
The last row is the one that matters. A chat gives an answer that is never afterwards compared with anything, so it has no error rate and cannot have one. A process that goes and asks five people ends up with a number for how wrong it was — and a tool willing to record that number is making a different kind of claim than one that is merely confident.
Our own limits, stated
A page that only listed the competition’s problems would have the same defect it describes. So here are ours, in the same words we would use if you asked us privately.
1The conversation itself does not measure demand — nothing said in fifteen minutes becomes a customer. What measures it is the step after: a landing page published at its own address that collects real signups, and a survey that refuses to score anything until at least five real people have answered. Both sit behind a credit, and both are work you still have to do.
2The synthetic focus group is synthetic. Ten to twenty personas built from public writing are a rehearsal of objections, not a sample of your market, and we say so on the page that generates them.
3Publishing a landing page is not the same as having visitors. We generate the page and record who signs up; getting people in front of it is still your problem, and a demand test with nobody in it measures nothing.
4The free tier gives the summary. The full report set is generated before payment and stays locked until a credit is spent, which we would rather state here than have you discover it.
5A verdict is a decision aid. It is computed from thresholds rather than composed by a model, precisely so that it cannot be talked into a nicer answer, but the decision stays yours.
Be asked, not scored
Fifteen minutes of questions that push back where an answer is thin, then the research most founders skip. You leave with a verdict and, more usefully, with the one thing worth finding out this week.
Start free →15 min · No credit card · summary free