Honestly: I don't have the data yet. The directory only just got rebuilt this
way and it goes live properly tomorrow, so anything I told you about user
behaviour right now would be me guessing with confidence I haven't earned.
My hypothesis is that you're right — most people will judge from the output and
the description, and the stored run will go mostly unopened. But I think that's
fine, because I've come to see the evidence as doing its real work on the supply
side rather than the demand side. Its main job isn't to convince a reader. It's
that a prompt which doesn't run, or runs and produces nothing useful, can't get
published in the first place. The badge is a byproduct of a gate.
That reframes the hash-binding too. It's not there so users can verify a hash —
nobody is going to do that. It's there so I can't quietly edit a prompt after it
passed and keep the badge. It constrains me more than it informs them.
What I'll actually measure after launch: whether people expand the evidence
drawer at all, and whether prompts with visible output get copied at a different
rate than ones where the drawer stays shut. If the answer is "nobody opens it and
copy rates are identical," that's a real finding and I'd rather publish it than
bury it.
Curious what would move you personally — would seeing the run change whether you
trusted a prompt, or would you just try it and judge the result yourself?
That’s useful context. I’d rather continue this conversation privately than go deeper here. What’s the best email to reach you on?
About
Every prompt site published strings nobody had run, mine included. So our 22K-member community's directory has one rule: nothing publishes until it's run in a sandbox, output visible. AI agents review, not humans.
3 Comments
The hash-bound evidence is the interesting part here.
Curious whether users actually treat the run evidence as a trust signal, or whether they still judge a prompt mainly from the output and description.