1
6 Comments

Everyone blames onboarding for churn. I read 2,797 angry reviews — half of them are about billing.

I crawl an app store for a living (Shopify's), which means I have a pile of reviews sitting in a database. Every time someone in a founder forum asks why users leave, the answer is onboarding. Time to first value. The first ten minutes. It's unanimous enough that nobody checks it.

I had 60,278 reviews across 906 apps, so I checked it.

Of the 2,797 one- and two-star reviews, 18.7% mention billing — charged after uninstalling, a charge they didn't expect, a refund they couldn't get. Setup and usability: 3.6%. Among the reviews that name a reason at all, billing is 48% and onboarding is 9%. The thing everyone optimises is the fifth most common complaint.

The control matters here, so: billing language is 14x more common in a one-star review than in a five-star one. It isn't just that people talk about money a lot.

The second finding is the one that changed how I think about my own product.

The platform publishes how long each reviewer had the app installed. Median tenure of a five-star review: 30 days. Median one-star: 90 days. Two-star: 120. The angriest users are the longest-tenured ones, not the confused new ones.

And 25.7% of all five-star reviews are written within a day of installing — before the person can possibly know whether the thing works. A quarter of your social proof is a review of your install flow.

Those two findings fit together rather than fighting. Someone confused on day one uninstalls quietly during the trial and never writes anything. Someone who has been paying for three months has a bill to be angry about.

What this cannot tell you, and it's load-bearing: reviews are not uninstalls. Silent day-one churn could be enormous and completely invisible here. If the onboarding theory is right, its evidence lives somewhere I can't see. What I measured is why the people who stayed long enough to get angry got angry. 61% of the negative reviews name no reason at all, and the themes are keyword matches, not a human coding 2,797 reviews — which is exactly why I report the five-star rate beside every theme.

Three things I'd do with this:

  • Audit the charge, not the funnel. What happens when someone cancels mid-cycle, what happens at a plan boundary, how a refund request gets handled. That's where half the written complaints live, and it's policy rather than engineering — unusually cheap to fix.
  • Stop reading day-one reviews as validation. A rising rating can sit right on top of churn you can't see.
  • Point your surveys at month three. The verdict forms around 90 days; almost everyone's check-in email is aimed at week one, where the opinions are cheapest.

Full method, charts and the remaining caveats: https://bestappify.com/blog/onboarding-is-not-why-they-leave-your-shopify-app-we-read-2-797-bad-reviews-47

Curious whether this holds outside app stores — if you've got cancellation reasons tagged, does billing beat onboarding for you too?

on August 26, 2026
  1. 1

    This is the best kind of data post, because your own load-bearing caveat is a bigger finding than your headline, and you've stated it without fully cashing it in. "Reviews are not uninstalls, silent day-one churn is invisible here" isn't a limitation, it's the actual mechanism, and it indicts how the whole industry measures churn.

    Watch what your two findings do together: complaint volume is inversely correlated with how early the failure happened. The confused day-one user uninstalls silently and writes nothing. The angry 90-day payer has a bill and a keyboard. So every voluntary-feedback channel, reviews, testimonials, support tickets, structurally undercounts early failure and overcounts late failure, by construction. Onboarding doesn't look like a small problem because it is one, it looks small because its victims don't file reports. That's survivorship bias, and the entire industry is reading it as a finding.

    Which means the reason "everyone blames onboarding" and the reason "your data says billing" can both be true at once, and neither is measured right. Billing is the loudest documented complaint. Onboarding might be the largest silent one. Your data can see the first and is blind to the second, exactly as you said, and the trap is concluding from "billing dominates the reviews" that onboarding is overrated, when your instrument can't see onboarding's body count at all. Argument from silence, mistaken for argument from absence.

    So the missing tool isn't "surveys at month three," it's measuring the silent ones behaviorally: exit timing relative to first-value timing. Uninstalled before the first successful action versus after three paid months are two different products with two different fixes, and only one of them ever writes a review.

    Do you have install-to-uninstall tenure on the apps themselves, not just the reviewers? That distribution is the one dataset that could actually see the churn your reviews can't.

    1. 1

      You're right that this is the mechanism, not just a caveat. What we can see is reviewer tenure, not install tenure — and those are only the same distribution if silent uninstalls and vocal complaints happen at the same rate over time, which they obviously don't.

      To answer directly: no, we don't have install-to-uninstall tenure on the apps themselves. We don't have visibility into every install and uninstall, only into who wrote a review. What we have is 2,797 one and two star reviews, and inside that set the median 1-star reviewer had paid for 90 days while a quarter of 5-star reviewers wrote within a day of installing. That's consistent with your read — angry, billed people write reviews, confused, gone people don't — but it can't separate "onboarding is a small problem" from "onboarding is invisible to this instrument." You're right that we can't rule out the second one, and the post shouldn't have implied we could.

      (I work on BestAppify, that's where the numbers come from)

      1. 1

        Your concession is the right call, and it makes the post better, not weaker — but don't fully write off seeing the silent layer, because you're closer to it than "we only have reviewers" suggests. You can't see an individual silent uninstall. You can see the shadow it casts across your 906 apps.

        Here's the move: silent early churn leaves a signature in the review distribution itself. An app bleeding day-one users produces lots of installs, few reviews, and the reviews it does get skew long-tenure, because the short-tenure people are gone before they'd ever write one. So reviews-per-install plus the tenure shape becomes a proxy for the thing you can't measure directly. An app with a healthy onboarding funnel and an app hemorrhaging silently will have differently shaped review-tenure curves even if their star averages match.

        You already have the one dataset that makes this work: 906 apps to compare against each other. You can't measure silent churn in absolute terms, but you can rank apps by how much their review distribution looks like silent-early-churn versus late-billing-anger. The billing-heavy, long-tenure-skewed apps are one failure mode; sparse-reviews-relative-to-category, tenure-skewed apps are the other. Same corpus, second signal.

        If you built that comparative view, the interesting test: do the apps with the worst reviews-per-install ratio also have the fewest billing complaints? Because if onboarding-death and billing-anger are different apps, not different reviewers, that would be the first real evidence for the silent theory your review data alone can't reach.

        1. 1

          This is a genuinely good idea and I want to build it, but I have to be straight about what's missing: we don't have install counts. The 906-app dataset is reviews only, so "reviews-per-install" isn't something I can compute today — I'd need Shopify to expose install numbers, which it doesn't, and estimating them would put a number in front of you that isn't real data.

          What I do have is the tenure field on every review, which is half of what you're describing. So the version I could actually run is narrower than yours: for each app, look at the tenure distribution of its reviews and see whether it skews long (few short-tenure reviewers) versus flat. Cross that against billing-complaint share per app. Your prediction is that those two things should anti-correlate — apps that are heavy on billing complaints shouldn't also show the long-tenure skew that silent early churn would produce.

          That's a real, checkable test on data I already have, just without the install-base normalization that would make it clean. I'll run it. If it comes back suggesting two distinct failure modes rather than noise, that's worth its own post, and I'll credit the framing to this thread.

          (I work on BestAppify — that's where the 2,797-review dataset comes from)

  2. 1

    This is the measurement clarity win. The founder community had a consensus: onboarding is the problem. But your 60,278 reviews measured the actual distribution of complaints - half the frustration is billing, not day-one friction.

    Two things matter here:

    1. You measured what customers actually complain about, not what you assumed they'd complain about
    2. The complaints revealed two different user populations with different failures (trial users vs paying customers)

    Most teams optimize for the consensus problem. You measured it and discovered the consensus was half wrong. Now your roadmap can route to the actual pain (billing for retention, onboarding for activation) instead of fighting the wrong problem with the wrong fix.

    This is the core of why measurement precision matters - billing complaints look like they're about billing. They're actually about "I paid and it didn't work." Completely different fix than "I installed it and didn't understand it." Same word (churn), two separate problems, measured apart.

    1. 1

      One correction to the two-populations read, because the data's actually a bit stranger than trial vs. paying. Both groups here are already paying customers. The median one-star reviewer had been on the app for 90 days before they wrote it, and the median five-star reviewer for 30 days, with a quarter of the five-star reviews posted within a day of installing.

      So it's less "trial friction vs. paid friction" and more: people who stick around, keep getting billed, then hit a wall (refund refused, charged after uninstalling) and write furious after months of loyalty. Versus people who are still in the honeymoon window and rate on vibes. That's actually a harder routing problem than activation vs. retention, because the angriest users aren't new — they're your longest-tenured ones.

      2,797 reviews, 906 apps, so decent spread, but it's Shopify app reviews specifically, so I wouldn't assume the same 90-day pattern holds outside that market. (I work on BestAppify, that's where the numbers come from.)