GUIDE
Selecting a playtest cohort
Who is in the lobby determines what the session can tell you. This is how to specify a multiplayer cohort so that the answer you get is about your build rather than about who happened to be free on Thursday.
Derive the cohort from the decision
The first question is not “how many players”. It is “who would have to struggle with this for us to change our minds”.
If the decision is whether new players survive the first match, a cohort of people who already love the genre cannot produce a negative result. Not because the build is good, but because the instrument cannot detect the problem. If the decision is whether a high-level counterplay loop works, ten curious newcomers will never reach the situation you need to observe.
Write the cohort as the answer to that question, then check it against what you can actually schedule. Where the two disagree, change the question or accept a narrower scope. Do not quietly run a cohort that cannot answer the question you wrote down.
Experience bands
Skill is not one dimension, and “experienced players” is not a specification. Four bands cover most multiplayer studies:
- New to the genre
- Has not meaningfully played this kind of game. The only band that can tell you what your onboarding actually does. Also the band whose participants most often blame themselves for a design problem, so their behaviour is more informative than their self-report.
- Genre-familiar
- Knows the conventions, does not know your game. Usually the closest match to the player a commercial decision is really about, and the band most teams under-recruit.
- Experienced
- Plays this genre regularly and competently. Good at articulating comparisons against other titles, which is useful and also a bias: they will tell you how your game differs from the market leader whether or not you asked.
- Competitive expert
- Plays at a high level. Essential for questions about depth, counterplay and balance ceilings. Actively misleading for questions about accessibility, because they route around friction without noticing it.
Specify a target count per band, not an average. “Twelve players, mixed experience” is not a cohort; it is what you will end up with by accident.
Include exploration seats deliberately: a small number of participants outside your expected profile. They are how you find out that your assumed audience is wrong, which is a finding no perfectly-targeted cohort will ever produce.
Roles and format
Specify the multiplayer format first (duel, co-op, four-player squad, 5v5, large lobby) because it fixes the arithmetic of everything else. Then specify roles if your game has them: entry, anchor, support, shot-caller, flexible.
Two practical points. First, a lobby of strangers left to self-select will produce a role distribution nothing like your live population; if role mix matters to the question, assign it. Second, shot-callers are scarce and disproportionately shape a session. One confident voice can turn an uncoordinated group into a functioning team, which is either your finding or your confound depending on what you were testing.
If the study is about first-time players, treat that as its own cohort cell with its own count, not as a footnote on the experience mix.
Hardware tiers
Split the roster across low, mid and high tiers in proportions that resemble your actual player base, not your studio.
The common mistake is recruiting on enthusiasm, which selects for people with good machines, which quietly removes the performance problems from your test. If a third of your players will run the game at 45 fps in a busy fight, a third of your cohort should too, or you will not learn what that feels like until reviews tell you.
Mid-tier is usually the hardest seat to fill and the most informative. Plan for it early rather than filling it with whoever is left.
Region, route and availability
Region does three separate jobs, and it is worth keeping them apart: it determines network conditions, it determines what time the session can run, and it carries cultural and market context about how the genre is played.
Specify the region and the named server route the session is played against, and require recent per-participant measurements to it. “Australia” is not a route; Sydney to a Sydney host and Perth to a Sydney host are meaningfully different seats.
Availability belongs in the cohort specification, not in the scheduling afterthought. A cohort defined without a time window is a cohort you cannot assemble. Store availability against a real time zone and confirm every offer in the participant’s local time.
Language matters for the same reason voice does: a session where two participants cannot follow the callouts is testing something other than your design.
Evidence quality: self-reported versus measured
Every attribute you select on has a provenance, and the provenance decides how much weight it can carry:
- Self-reported
- What the participant typed. Fine for genre familiarity and preferences. Weak for hardware, where people are sincere and frequently wrong, and weak for skill, where self-assessment is famously unreliable in both directions.
- Browser-measured
- Values captured by an instrumented check the participant runs themselves: GPU string, core count, approximate memory, display, a timed render loop, and round-trip and jitter to a named route. Approximate, but not a self-assessment.
- Operator-verified
- Confirmed directly by a researcher. Strongest, and the most expensive to obtain.
- Session-observed
- What the participant actually did in a previous session. The best predictor of what they will do in the next one, and only available after they have attended one.
Set an evidence policy per study, deliberately: verified only, browser-measured or verified, or self-reported allowed. Stricter policies cost fill rate and lengthen recruitment. Looser policies cost certainty about the conditions your findings were produced under. There is no default that is right for every study, which is why it should be a recorded choice rather than an implicit one.
Treat measured evidence as perishable. Route quality changes with an ISP or a house move; drivers change monthly; a GPU is stable for a year. Expiry windows should differ by evidence type, and an expired capture should read as expired rather than be silently reused.
Reliability, and why rate must not rank
Past attendance and readiness behaviour is a legitimate selection input: someone who has shown up ready twice is a better bet for a session that cannot be re-run. Participants with no history need exploration seats, or the panel calcifies into whoever was early.
What must never rank is price. The moment a cheaper participant is preferred, you have built a reverse auction: participants learn to underbid, the cohort skews towards whoever most needs the money, and your selection criteria stop being about research quality. Use a fixed posted reward, make any premium explicit, and check budget feasibility after selection rather than during it.
Size, and the standby ratio
Size the cohort from the format, not from a sense of statistical comfort. A 5v5 study needs ten primaries; running eight because two people are unavailable is a different study. Below roughly eight participants a synchronized session is usually the wrong instrument and a focused, tailored method is better.
Then add standbys on top: a meaningful fraction of the roster, not one token spare, and pay them whether or not they are called in. The right ratio depends on your lead time, how well you know the participants, and how catastrophic a short lobby is for the protocol.
Resist scaling up for statistical confidence. At the sizes a moderated multiplayer study can realistically schedule, more participants buys you more observations and more coordination risk, not significance. If you need distributions, that is a telemetry question, not a cohort question.
Record the cohort you got, not the one you specified
The specified cohort and the attended cohort are rarely identical, and the difference qualifies every finding. Record both, plus the substitutions, the seats that went unfilled and why, and the evidence provenance behind each attribute you selected on.
When someone asks in three months whether a finding still holds for the current build, that record is what makes the question answerable instead of a matter of memory.
Playtest Live