How many testers you need before you can ship
Pattern over sample size theater. Two or three cold runs on the same brief usually beat a dozen vague sessions — then stop, fix, or ship.
Indie developers ask for a number: how many testers before I can ship? The honest answer is not “always five” or “ten for statistical significance.” It is: enough cold attempts on the same brief to see a pattern — then act.
Sample size theater wastes money. You buy twelve slots, get twelve different missions in practice, and still cannot approve or reject anything. Structured QA flips that: one scary path, the same form for every submission, comparable recordings. The question becomes what did independent people do on this path? not how big was the panel?
Start with one mission, not a headcount
Before you count testers, name the ship decision:
- Ship — the scary path completes for cold people.
- Hold — blocked, but you need one more data point to be sure.
- Fix — two or more people hit the same wall; that wall is the release blocker.
Pick the single path that would embarrass you in support if it fails: checkout, invite, first project, plugin activation + settings save. Write it as three to seven observable steps with a clear success state. That brief is the unit of measurement — not “testers.”
What two or three testers actually tell you
On a tight, high-signal brief:
- All finish — strong signal to ship that path (not “the whole product is perfect”).
- All blocked at the same step — strong signal to fix before ship; you have a reproducible stall.
- Split results — often environment or brief ambiguity, not “users are random.” Tighten credentials, staging parity, or step wording; re-run the same brief.
- One blocked, two finished — investigate the outlier recording. If it is access friction (bad login, 2FA), fix access and re-run. If it is a real edge case, decide whether that role matters for this release.
Three comparable submissions on one mission beat ten diary entries about “overall impressions.” You are looking for convergence, not a confidence interval.
When to add a fourth or fifth slot
Add testers when you still cannot decide — not because bigger feels safer:
- After a fix. You patched the stall; re-run the identical brief with fresh cold people. That is a new round, not padding the old one.
- High-stakes money path. Billing, payouts, or permission changes that touch revenue — a second trio on the same brief after a clean first trio is reasonable.
- Role matrix. Admin works; editor does not. That is a different brief, not “more testers on the admin path.”
If five people on the same brief still produce chaos — contradictory forms, recordings that do not match answers — the brief is broken. Fix the mission before you buy more slots.
When to stop and ship
Stop widening the sample when:
- The scary path completes for at least two independent cold attempts with recordings you would approve.
- No open blocker remains on that path — or you explicitly accept a known edge case for this release.
- CI and smoke on the artifact you will ship are already green (scripts first; humans second).
Ship does not mean “nobody will ever complain.” It means the path you were afraid of is no longer the obvious failure — on evidence you can point to, not hope.
What does not count toward “enough”
- Your own walkthrough — you are not cold.
- Twelve testers on twelve different informal tasks.
- An open beta with no form — noise, not settlement.
- Exploratory sessions with no success criteria — you cannot approve or reject them.
- Buying more slots instead of fixing a clear pattern at step 4.
If you cannot reject the evidence, you are not gating. You are hoping with a larger bill.
A practical default for indie teams
For most releases:
- One brief, one scary path, 2–3 slots.
- Read submissions together: same stall? fix and re-run the same brief.
- Ship when the path is clean — or hold with a named blocker, not vague unease.
Tomorrow’s release gets a new mission. You do not need a permanent panel — you need a repeatable habit.
On QATested
QA Testing on QATested is built for small, decisive rounds: post one brief, set a few slots, get recordings + structured answers, approve or reject, unused slots refund. How to post: post a Structured QA test.
Related: how indie developers run a paid QA test · ship checklist for a one-person team · brief, high-signal tests · structured QA vs an open beta.