1
0 Comments

An AI-built app passed my test — generated code was not the pass condition

Most AI app demos stop when the code appears or the preview renders. I wanted a stricter definition of success.

I tested Modelence with one synthetic operations-manager persona and one prompt: build an employee timesheet.

The evaluator took 21 recorded steps, entered the live app, clocked in, clocked out, and left a visible 21-second entry.

That final state is why I marked the run as completed. Generated code alone would not have counted. A button click alone would not have counted. The target user had to complete the intended outcome and leave something inspectable behind.

The report also kept the limits visible:

  • One synthetic persona
  • One simple task
  • One friction: no clear one-click publish path before the live sandbox
  • No assessment of code quality, security, production readiness, adoption, or persistence after the observed session

I’m using this positive case to test a $29 August prelaunch round for founders who want one core product flow exercised before launch.

If you build AI products, what would you accept as a real pass: generated output, a working preview, or a completed user outcome in the live product?

on August 6, 2026