A validation prompt that argues back
Ask a chatbot whether your idea is good and it will tell you it is. Not because the idea is good, but because you asked a question with an agreeable answer.
The prompts below remove that path: the model has to argue the case against your idea first, name the competitors you forgot, mark which numbers it is guessing, and finish with a verdict it must defend.
Prompt 1: the validation pass
Works in ChatGPT, Claude and Gemini. Turn web search on if you have it, and fill the two blocks in brackets.
Prompt 2: find the complaints
Run this second. It looks for how people describe the problem in their own words, searching for complaints rather than recommendations.
Now test the answer you got
Do not take our word for the limits. Three checks, five minutes, on your own idea and your own answer.
Ask the same idea again in a fresh chat
Same prompt, same idea, new window. Compare the two verdicts. They will not match, and usually not by a little: 6/10 and 8/10 on the same paragraph is normal. Nothing in a prompt anchors the score to anything measured, so what you got was a mood, formatted as a rubric. If the second answer is better news than the first, notice how much you want to believe that one.
Try to click through to one number
Take the market size it gave you and ask where it came from. You will usually get a plausible-looking attribution to a research firm and no link, because the figure came out of training data rather than a document. Now imagine that number inside a deck, in front of someone who checks. A confident paragraph reads exactly the same whether the source exists or not.
Ask it to name five competitors, then check them
Search each one. Expect at least one that does not exist, and at least one real company that shut down or pivoted a year ago. Neither is a bug you can prompt away: one model in one pass has no way to catch its own invention, and no reason to hesitate before writing it.
None of this makes the prompt useless. It makes it a first pass — a way to find the questions worth answering, not the answers themselves. The failure mode is not that the output is bad. It is that the output is convincing, and you stop there.
What closing those gaps requires
Each of the three checks fails for a structural reason, and none of them is fixable by writing a better prompt.
| The gap | What it takes to close it |
|---|---|
| Different verdicts each run | The verdict is computed, not written. Market size, growth, business model, problem clarity and audience each score against fixed thresholds, and the verdict is a threshold on the sum. Run it twice and the number moves only if the evidence moved. |
| Numbers without sources | Six research stages run with live web search, and every claim carries the link it came from. Where a figure is an estimate, the report says so. |
| Invented competitors | Key claims go through three independent models from three vendors plus web evidence, and when they disagree the disagreement is recorded by name rather than averaged into a smooth sentence. |
And the input differs before any of that starts: instead of a paragraph you typed, a 15-minute conversation that asks follow-up questions and pushes when an answer is thin. Measured on our own projects, that brings ten times more context into the research — median 1,616 characters of your own words against 163 from a single field. The prompt above is limited by what you thought to write down. The interview is not.
Run the whole thing instead
Same task, different machine underneath: a voice interview, six research stages with live search, three models checking key claims, and a computed verdict. Three projects free, no card.
Start free — 3 projects →15 min · No credit card · up to 20 reports