
Most AI app demos stop when the code appears or the preview renders. I wanted a stricter definition of success.
I tested Modelence with one synthetic operations-manager persona and one prompt: build an employee timesheet.
The evaluator took 21 recorded steps, entered the live app, clocked in, clocked out, and left a visible 21-second entry.
That final state is why I marked the run as completed. Generated code alone would not have counted. A button click alone would not have counted. The target user had to complete the intended outcome and leave something inspectable behind.
The report also kept the limits visible:
I’m using this positive case to test a $29 August prelaunch round for founders who want one core product flow exercised before launch.
If you build AI products, what would you accept as a real pass: generated output, a working preview, or a completed user outcome in the live product?