
Two months ago, I added a new rating system to my business idea discovery tool (https://1mil.app/ih). It asks, "Can one person win this niche?" It weighs incumbent strength, pricing pressure, and whether a newcomer has any moat available at all.
Since then, the tool has scored 982 new opportunities.
Six ever crossed the "Winnable" line.
I checked whether I could trust even those six.
Five of those six were scored by the first version of the rating. This month, I ran a test-retest check on that version: score the same ideas twice, correlate the ranks. The correlation came back at -0.05, which is nearly zero. So those five "winnables" were coin flips.
The rebuilt rating system scores each idea from multiple samples and prints the agreement count on the card. When the samples disagree completely, the card says "Signals disagreed" instead of showing a meaningless number.
Under the rebuilt system, exactly one idea earned "Winnable" and kept it when the samples were compared: a narrow, unglamorous workflow API in a niche nobody tweets about.
One out of 982... 😬
Three things to conclude:
Full post-mortem: https://1mil.app/learn/does-ai-idea-scoring-work
What would convince you that a "this niche is winnable" claim was real?
The test-retest result is the most interesting part. If the same opportunity can receive radically different scores, the real product isn't the score itself but making the evidence and disagreement visible enough to judge the conclusion.
50 shades of winnability becomes the real product.
Joking aside, that is close to where it ended up. Every card now shows the agreement count, and when the samples disagree completely it says "Signals disagreed" instead of a number.
What would "visible enough to judge" look like to you, in card form?
That's the part I'm least sure I've solved.
The “Signals disagreed” treatment is interesting. It makes the uncertainty visible without pretending the score is more precise than the underlying signal.
The retest is necessary, but it only shows repeatability. It doesn’t show that “winnable” predicts anything.
I’d define the outcome first: a specific revenue or retention threshold within a fixed period. Then backtest the scorer on historical niches whose later outcomes were hidden from it. The useful result would be calibration by score band: how often did ideas in each band actually cross the threshold?
Agreement across samples reduces noise. It doesn’t validate the claim.
And I did run one test in your direction: known bootstrapped winners, scored through the same pipeline, came back between 1.0 and 5.9. It does not cleanly separate winners. That's published in the site FAQ.
Your backtest has a structural problem here though: the scorer researches the live web, so I can't hide a niche's later outcome from it. The evidence source contains the ending.
The prospective version I can start today: fix a threshold up front, log score bands at scan time, publish the calibration table once enough time has passed. Where would you set the threshold and the window? Your spec is probably better than mine.
I’d use a deliberately modest threshold: three paying customers from the target niche within 90 days of first outreach. That tests whether the niche offers an accessible path, not whether it can become a large business.
The denominator matters as much as the threshold. Count only ideas someone actually pursued past a minimum effort floor, and publish conversion by score band. Otherwise unattempted ideas become false failures and builder execution gets mixed into the market score.
Kesha, you're right about the distinction. The retest shows the rating agrees with itself, nothing more. I went after repeatability first because without it any validity test would have been meaningless, not because repeatability proves the label works.
On what the label actually claims: it describes the market as it exists today (incumbent strength, pricing pressure, whether any moat is open to a newcomer). It doesn't forecast the builder's outcome. The builder is the experiment. That's also why I don't present it as a prediction.