I went into yesterday's Shanghai Indie Hacker meetup thinking Menso's strongest value was downstream analysis: give an AI user a task, study its path, and turn the friction into a UX report.
Marc Lou gave me a more useful product order.
First: did the AI user complete the task?
Second: if it failed, where exactly did the task become impossible?
Only then should Menso explain the broader UX friction.
That reframes the product from “an AI that comments on a journey” to a system that verifies an outcome and finds the first interruption point.
The distinction matters because many product failures look healthy from the surface. A button can accept a click while the object is never created. A funnel can reach its final screen while the intended state is never persisted. A report that starts with general UX observations can miss the more important fact: the task did not actually succeed.
The core sequence I am working toward is now:
Menso already records the browser run and produces a prioritized report. My next product decision is how much of the outcome verification should be explicit and structured before the deeper analysis begins.
For anyone building agent tests or product QA: how do you define “task completed”? Is reaching the expected screen enough, or do you require evidence that the intended object or state exists?
I am also testing a simple paid offer this month: $29 for one AI-user test of a core flow, with replay, screenshots, and a prioritized report delivered within 24 hours.