Skip to main content
← All founder guides
Straight answer · what AI can and cannot settle · 8 min read

Can AI validate my startup idea?

We sell one of these tools, so start with the part that argues against us: a model reading your description is scoring your writing, not your market. Two founders with the same business get different verdicts depending on who writes better, and either of them can raise the score by editing the pitch instead of changing the company.

That objection is correct. It is also not the whole line. Below is where the line actually falls, what moves it, and what nothing moves.

Short answer

Can AI validate my startup idea?

Partly, and the part matters. An AI can check whether your reasoning holds together, who already sells this, whether your market arithmetic survives one question, and whether your customer questions are leading. It cannot establish demand, because demand is a fact about other people's behaviour.

Where the line falls

  1. An AI can settle four things from text alone: internal contradictions, who the competitors and substitutes are, whether the market arithmetic is defensible, and whether your planned questions are leading
  2. It cannot settle demand. The only instruments for demand are a payment, a deposit, a signed commitment, or an alternative someone cancelled
  3. A confident verdict on demand derived from a paragraph is scoring your writing: edit the pitch and the number moves while the business does not
  4. A 15-minute interview changes the input, not the laws — it captures what you did not think to write down, and exposes the gap between a claim and its evidence
  5. The useful output of any validator is not a score but an ordered list of what to go find out, cheapest question first
  6. No prompt can add a step that does not exist: being asked rather than typing, three models checking one another, a landing page real people sign up on, a survey that needs five real answers, and a record of how far the earlier prediction missed
  7. 42% of startups fail for lack of market need (CB Insights) — which is a reason to go ask people, not a reason to trust a score

What text can settle, and what it cannot

The split is not about how good the model is. It is about whether the answer exists in language at all, or only in someone else’s behaviour.

QuestionWho can answer itWhy
Whether your reasoning holds togetherAn AI canContradictions between your market size, your price and your channel are visible in the text itself. So is a plan whose first hundred customers arrive through "social media, SEO and word of mouth".
Who already sells this, including the boring substitutesAn AI canCompetitors are public. So are spreadsheets, agencies and doing nothing, which are the real incumbents and the ones founders forget to count.
Whether your market arithmetic is defensibleAn AI canA top-down number derived from a global market figure falls apart under one question about who is actually addressable. That question can be asked mechanically.
Whether your customer questions are leadingAn AI canA question that pitches, asks for an opinion, or asks about the future instead of the past is recognisable from its wording alone. That is a language problem, and language is what models are good at.
Whether anyone will payNothing can, except paymentDemand is a fact about other people. The instruments are a payment, a deposit, a signed commitment, or an alternative someone cancelled. Enthusiasm about a hypothetical is not one of them.
Whether you will still care in a yearNothing canThe most common way a small company dies is the founder losing interest, and no market analysis has ever predicted it.
Whether your channel works at your costOnly running itChannel economics are specific to your product, your price and your moment. A general estimate here is astrology with decimals.

Why a score is the wrong output

A single number invites one question — is it high enough — and that question has no useful answer. Worse, it is editable: rewrite the description more confidently and the number rises while nothing about the business has changed. Any output you can improve without leaving your desk is measuring the desk.

The output that survives contact with reality is an ordered list of unknowns, where each item says what would count as an answer and roughly what finding out costs. Then validation stops being a verdict you receive and becomes a week you spend.

This is why every tool on this site shows how much to trust its own result: a range instead of a point, the share of an answer that was invented, the worst case next to the average. A tool that hides its uncertainty is not being confident, it is being quiet.

What a 15-minute interview changes

It changes the input, not the laws. A description contains what you thought to write down. Being asked follow-up questions for fifteen minutes contains what you did not, and a thin answer invites another question where a paragraph simply sits there.

Concretely, in our case: three consultants ask in turn, the conversation runs in 21 languages, competitors are pulled live rather than recalled, key claims are cross-checked by three independent models, and the verdict is computed from thresholds instead of composed by a model — so it cannot be talked into a kinder answer. The result is up to 20 reports, of which the summary is free.

And what it does not change: the conversation itself still does not measure demand. It tells you which of your assumptions is load-bearing and which question to go answer first. Measuring demand is the next step, and it is a step rather than a smarter analysis — which is the whole point of the section below.

What actually closes the loop

Most of the writing on this subject argues that a chatbot is too agreeable, and cites good research to prove it. The argument is true and it is also easy to answer: if the problem were the model’s manners, a better prompt would fix it, which is why prompt packs sell.

So here is the harder version. No prompt can add a step that does not exist. The difference between a conversation and a validation is how many of the steps touch the world outside the conversation.

1Being asked, not typing

The input to a chat is what you chose to write down. Fifteen minutes of adaptive follow-ups is what you did not, and a thin answer draws another question instead of passing.

2Checked by other models

Key claims are cross-checked by three independent models. One model agreeing with itself is not verification, however the prompt is worded.

3A verdict computed, not composed

The GO / NO-GO comes from thresholds rather than from text generation, which is why it cannot be talked into a kinder answer by a more confident description.

4A landing page that real people see

The brand consultant generates a landing page and it is published at its own address, where real visitors sign up and those signups are recorded. That is demand, not an opinion about demand.

5A survey real people answer

Questions are generated from your own hypotheses and sent out as a public link. Nothing is scored until at least five real people have answered.

6Our own error, written down

When those answers come back, the system computes how far its earlier prediction was from what people actually said, and stores that gap. A chat answer is never compared with anything, ever.

The last row is the one that matters. A chat gives an answer that is never afterwards compared with anything, so it has no error rate and cannot have one. A process that goes and asks five people ends up with a number for how wrong it was — and a tool willing to record that number is making a different kind of claim than one that is merely confident.

Our own limits, stated

A page that only listed the competition’s problems would have the same defect it describes. So here are ours, in the same words we would use if you asked us privately.

1The conversation itself does not measure demand — nothing said in fifteen minutes becomes a customer. What measures it is the step after: a landing page published at its own address that collects real signups, and a survey that refuses to score anything until at least five real people have answered. Both sit behind a credit, and both are work you still have to do.

2The synthetic focus group is synthetic. Ten to twenty personas built from public writing are a rehearsal of objections, not a sample of your market, and we say so on the page that generates them.

3Publishing a landing page is not the same as having visitors. We generate the page and record who signs up; getting people in front of it is still your problem, and a demand test with nobody in it measures nothing.

4The free tier gives the summary. The full report set is generated before payment and stays locked until a credit is spent, which we would rather state here than have you discover it.

5A verdict is a decision aid. It is computed from thresholds rather than composed by a model, precisely so that it cannot be talked into a nicer answer, but the decision stays yours.

Be asked, not scored

Fifteen minutes of questions that push back where an answer is thin, then the research most founders skip. You leave with a verdict and, more usefully, with the one thing worth finding out this week.

Start free →

15 min · No credit card · summary free

Frequently asked questions

Can AI validate my startup idea?+
Partly, and the part matters. An AI can establish whether your reasoning holds together, whether competitors exist and who they are, whether your market arithmetic is defensible, and whether the questions you plan to ask customers are leading. It cannot establish demand. Demand is a fact about other people's behaviour, and the only instruments that measure it are a payment, a deposit, a signed commitment, or a cancelled alternative. Any tool that returns a confident verdict on demand from a paragraph of text is scoring your writing, not your market.
Are AI idea validators accurate?+
Ask what they are being accurate about. Most score the description you submitted, which means two founders with identical businesses get different scores depending on how well they write, and the same founder can raise the score by editing the pitch rather than changing the business. That is not a measurement, it is a rewrite. The honest test of any validator is whether it can ever tell you something you did not put in, and whether it names the confidence of its own answer instead of returning a single number.
What can an AI validator never tell you?+
Four things. Whether anyone will pay, because that requires someone paying. Whether you will still want this in a year, which is the most common cause of death and no market analysis predicts it. Whether your particular channel will work at your particular cost. And what your customers will say in their own words, unless it actually goes and asks them. A validator that skips all four and still returns a verdict is confident about the easy parts.
What does a voice interview change?+
It changes the input, not the laws. A written description contains what you thought to write down; fifteen minutes of being asked follow-up questions contains what you did not. It also exposes the gap between a claim and the evidence behind it, because a thin answer invites another question and a paragraph does not. What it does not change: the conversation itself still does not measure demand. Demand is measured by the steps after it — a landing page published where real people can sign up, and a survey that scores nothing until at least five of them have answered, at which point the process also records how far its own earlier prediction missed. No wording of a prompt adds a step like that to a chat window.

More on this topic