been running aisa.to for a while now and we just crossed 400 assessments. here's the thing nobody expected:
people who use AI every day don't score much better than people who use it a few times a week. the average across all users is 52/100, and 62% score below what we'd call Proficient.
the actual differentiator? verification habits. people who systematically check AI output — not just "does this sound right" but actually spot-checking claims, asking the model to argue against itself, cross-referencing — those people score dramatically higher across every dimension.
safety was the weakest area at 45/100. most people don't think about data privacy, bias, or when NOT to use AI. that gap is going to matter a lot more as AI gets embedded deeper into workflows.
the counterintuitive takeaway for anyone building in the AI space: your power users aren't necessarily the ones using your product the most. they're the ones using it the most critically.
full data: https://aisa.to/state-of-ai-fluency
The verification habits finding resonates. Frequency builds familiarity but not necessarily judgment. We see the same pattern with Agent Builder — users who let agents run on autopilot get mediocre results; users who actually review what the agent does and push back tend to iterate toward much better output. We ended up making approval workflows the default rather than optional partly to nudge people toward that critical engagement habit.
The 45/100 safety score is the number I'd fixate on. Does the assessment show any correlation between safety awareness and verification habits, or are they independent failures?
"Your power users aren't the ones using it most, they're the ones using it most critically" is the sharpest line here, and a bigger finding than the post frames it as. If verification habits (not usage frequency) predict skill, that's not just a data point, it's your positioning handed to you by your own data.
Most AI-skill tools implicitly sell "use AI more / better." Your data says that's the wrong axis, the differentiator is critical verification, which almost nobody teaches and most tools actively discourage (they want more usage, not more skepticism). So Aisa's wedge isn't "assess AI fluency" broadly, it's "the tool that measures the one habit that separates skilled users from confident-but-wrong ones." That's a category of one. "AI fluency" is crowded and vague; "AI verification / critical-use skill" is specific, defensible, and backed by your own dataset.
The safety-at-45 finding compounds this and points at your B2B motion (which we discussed last time). Enterprises don't lose sleep over whether employees use AI enough, they lose sleep over employees using it wrong, unsafely, no verification, leaking data or shipping hallucinated output. That's a board-level anxiety with budget behind it. "We measure whether your team uses AI safely and critically" is a far easier enterprise sell than "AI fluency assessment," because it maps to a fear someone already has.
So the read: your data is telling you to reposition from fluency (nice-to-have, B2C-flavored) to verification-and-safety (must-have, B2B-flavored, budget attached). Same assessment, sharper frame, and it resolves the B2B/B2C fork from last time by pointing at B2B.
That "your data is handing you your positioning" move is what I spend my time on, part of the team building Hivemind, an AI strategy copilot. If you want to pressure-test the verification-first reframe: https://hivemind.myosin.xyz. Either way, this dataset is a stronger asset than one post, mine it for the wedge.
The part that surprised me wasn't the average score - it was that daily usage wasn't the strongest predictor.
It feels like a good reminder that repetition and improvement aren't the same thing. Using a tool more often builds familiarity, but regularly questioning its output is what builds judgment.
That distinction probably becomes even more important as AI starts blending into everyday workflows.