Skip to main content
← The journalLibraryLog inSign up free
GoNoGo JournalNº 07
Product27 March 20266 minGoNoGo Team

Reality Check, Synthetic A/B and AI Digest: What Happens After the Verdict

An AI verdict is a forecast. GoNoGo checks it: Reality Check puts it in front of real people with Mom Test questions and shows the gap, Synthetic A/B tests the details, AI Digest watches the market.

Reality Check, Synthetic A/B and AI Digest: What Happens After the Verdict — illustration for GoNoGo blog article

Every AI idea validator ends with a score. Many of those scores are a model's opinion, and models are known to agree with the person asking. Ours is computed: researched values scored against fixed thresholds, with no room to round up. But even an honest calculation rests on forecasts, market estimates and a synthetic panel among them. So the useful question is not "what did the verdict say?" but "how far off was it?"

GoNoGo keeps three tools running after the verdict. One of them exists only to measure our own error.

Reality Check: the verdict meets real people

Reality Check turns your project into a short survey, you send the link to people you consider your audience, and GoNoGo compares their answers with what the AI predicted.

  1. Questions from your hypothesesEvery question is tied to one hypothesis from your validation: a red flag, a green light, a persona, pricing or market.
  2. You share the linkRespondents see a neutral brief, need about two minutes, and give no name or email.
  3. The gap is measuredOnce responses arrive, each hypothesis is marked confirmed, rejected or inconclusive, and the reality score is set against the AI score.

Why friends can't flatter it

The obvious objection: people you know will be polite. That is why the questions follow The Mom Test. They ask about past behaviour and the current situation, never about your idea.

AsksNever asks
When did this problem last happen to you?Would you use a tool that solves it?
How do you deal with it today?Do you like this idea?
What do you currently spend on it?How much would you pay for this?

The survey generator is instructed to leave out hypothetical questions, two questions in one, leading questions and opinion questions; you review the result before anything is published. The brief respondents read starts with "Someone is building…", with no sales language. Politeness has nothing to hold on to: nobody is asked whether they like anything.

The first round has 8 to 12 questions in a fixed order: warm-up, the problem, current solutions, frequency and intensity, spending, one open question. Two or three screening questions use roles and contexts from your market, not generic "founder / developer" options.

What you get back

Analysis starts at 5 responses. The main number is simple: reality score minus AI score.

illustrative example, not a real project
Hypothesis: freelancers lose invoices every month → confirmed
Hypothesis: they already pay for a tool → rejected
Hypothesis: the pain is weekly, not quarterly → inconclusive
AI score 68
Reality score 57
Gap -11 (yellow zone: 10 to 20 points)
Direction the AI was optimistic
Next: round 2 asks only about the rejected and inconclusive hypotheses

A gap within 10 points is green, up to 20 is yellow, beyond that is red. The gap is written into the project and added to your reports rather than hidden. A second round of 6 to 10 questions goes only after what stayed unclear or was rejected; confirmed hypotheses are not asked again.

Reality Check also looks at your Synthetic A/B runs: did the option the panel preferred match what real people confirmed?

Note
Reality Check is part of a validated project. In the sidebar open Tools → Reality Check, pick your project, review the questions (you can edit, remove or add them) and publish the link.

Synthetic A/B: test the details before real people do

Once the idea itself is validated, the open questions are smaller: which price, which positioning, which feature set.

  • Compare 2 to 4 variants on one dimension: pricing, value proposition, positioning, feature set, or your own.
  • The judges are the same personas as in your focus group, built from real posts in communities, and each one remembers its earlier score.
  • The order of variants is shuffled for every persona and options carry neutral labels, so the first option gets no advantage.
  • Each persona gives a 1 to 10 score, a buying intent (would buy, might buy, would not buy), pros, cons and how sure it is.
Watch out
A synthetic A/B is a forecast, not a statistic. Use it to narrow options down, then put the leader in front of real people with Reality Check.

AI Digest: has anything changed since you validated?

Every morning GoNoGo writes search queries for your specific project and pulls news from the last seven days only: headlines, industry trends, competitor moves, funding and deals, and what it means for you. Every claim links to its source; older articles are dropped.

The digest is tied to Reality Check too. If real answers diverged from the AI forecast by more than 10 points, the digest looks specifically for why people answered differently than the model predicted.

One loop, not three tools

  1. ValidateA voice interview of about 15 minutes, six research phases, a calculated verdict
  2. TestSynthetic A/B on pricing, positioning and features with your own personas
  3. CheckReality Check: Mom Test questions to real people, the gap against the AI score
  4. WatchAI Digest: the last week of your market, and why reality disagreed

A validator that never checks its own forecast is asking you to trust it. This loop is built so you don't have to.

After reading

Validate your idea in a 15-minute talk

Pitch it to Alex out loud. He pushes back the way this journal does, then the research runs and you get a written GO, WAIT or NO-GO with the reasoning. Free to start.

Try Alex ↗3 questions · no signup
Next in the journal