GUIDE
Planning a multiplayer playtest
A single-player playtest degrades gracefully when things go wrong. A multiplayer playtest does not: it either assembles or it does not happen. This is what to plan for, in the order the problems actually arrive.
Start from the decision, not the session
Write down the decision the team will make differently depending on the result, before anything else. “Do we ship the new capture mode in the next build” is a decision. “Get feedback on the capture mode” is not.
The test is simple: if you cannot name an outcome that would change what you do on Monday, you are running a demo, not a study. Demos are fine, they are just much cheaper to run than a scheduled twelve-person session, and you should not pay session prices for one.
Everything downstream (cohort shape, session length, what you capture) is derived from that sentence. Teams that skip it end up with three hours of footage and an argument.
Scheduling: the constraint everything else bends around
You are not scheduling a session. You are scheduling the intersection of ten to twenty people’s evenings, and that intersection is much smaller than it looks.
- Pick the window from the cohort, not the team. If your players are in Sydney, Auckland and Singapore, there is roughly one usable weekday window and it is probably inconvenient for someone on your team. Choose it anyway; the cost of moving one developer’s evening is lower than the cost of a thinner cohort.
- Confirm in the participant’s local time. Every no-show post-mortem eventually finds someone who read a time in the wrong zone. Store an IANA zone, render local time, and never send a bare “7pm”.
- Give real lead time. A week or more for a scheduled evening. Booking inside a few days does not just risk no-shows, it selects for the people with nothing on, which is its own sampling bias.
- Budget the full commitment, not the play time. A “one hour session” is usually 20 minutes of setup and check-in, 45 to 60 of play, and 20 to 30 of debrief. Tell participants the real number before they accept, or you will lose people at the ninety-minute mark, mid-session.
- Re-confirm at T–24 and T–2. These are the two points where a session can still be saved. Treat them as scheduled operational steps, not as reminders.
No-shows and standbys
Plan for absence as a normal event. The failure mode is not that someone drops out; it is that you designed as if nobody would.
Book standbys as part of the roster, and pay them whether or not they are called in. An unpaid standby is a person who has been asked to keep an evening free for nothing, and they will treat the commitment accordingly. Paying standbys is not generosity, it is the only version of the arrangement that works twice.
Decide the promotion cut-off in advance, meaning the latest moment you will still swap a seat rather than run short, and decide what you do past it. Running a 5v5 with nine players is a different test; if you do it anyway, record that you did, because every finding from that session is scoped to those conditions.
Keep the swap on the record: who was replaced, when, and why. Six weeks later, when someone asks whether the coordination problem was real or an artefact of a missing player, that record is the only thing that answers it.
Regional latency
Latency is not a technical footnote in a multiplayer study, it is a variable that alters the behaviour you came to observe. Two players at 20 ms and 180 ms are not playing the same build.
- Name the server route and measure it. Which region the session is hosted in, and each participant’s actual round-trip and jitter to it. Measured recently, not measured once.
- Decide whether latency is controlled or observed. Either you hold it roughly constant so it is not confounding your finding, or you deliberately spread it and make it part of the question. Both are valid. Not choosing is not.
- Watch for the OCE and APAC trap. A cohort recruited across Australia, New Zealand and South-East Asia will have a wide route spread by default. If your question is about game feel, that spread is noise. If it is about whether the netcode holds for the region, it is the point.
- Treat route evidence as perishable. A route measurement from three months ago says nothing about tonight. Re-measure close to the session.
Hardware diversity
Studio machines are not player machines. A session run entirely on high-tier rigs will not surface the frame-rate cliff that a third of your players hit in the same fight.
Specify hardware tiers explicitly and fill them. The mid-tier seat is usually the hardest to fill and the most informative, which is a bad combination if you leave it to chance. Prefer measured hardware evidence over self-reported specs where it matters: people are honest and also frequently wrong about what is in their machine.
Then decide, before the session, what you do when a participant’s rig turns out to be different from what was recorded. Invalidate the seat, or keep it and label the finding? Either is defensible; deciding during the session is not.
Voice communication
Voice is the largest uncontrolled variable in most multiplayer playtests, and the one teams most often forget to specify.
- In-game voice and external voice are different tests. If players use an external client, you have not tested your in-game comms, and you may have accidentally tested whether your game works when comms are perfect.
- Strangers and teams behave differently. A cohort of people who have never met will under-communicate compared with a squad that plays together. Whichever you pick, it should match the population you are designing for.
- Voice is your best evidence source. Confusion, callouts and the moment someone says “wait, what just happened” are where findings come from. Get consent for recording explicitly, and separate the consent to record from the consent to be quoted.
- Check it before the session. A voice check belongs in the readiness list alongside the build launch. Discovering a broken microphone at start time costs the whole lobby, not one player.
Team composition
How you split the cohort into teams is a design decision that will shape your results.
Balanced teams give you a session that plays like a normal match. Deliberately unbalanced teams tell you what a losing player experiences, which is often the more commercially important question and almost never gets tested. Stacking every experienced player on one side produces a blowout, and blowouts produce disengagement data rather than the data you planned for.
Watch for one experienced player carrying a team through the exact friction you were trying to observe. If the study is about whether new players can figure something out, one expert in the lobby can hide the finding entirely. Either separate the bands, or record which team had which mix so the observation can be read correctly.
Assign roles explicitly if the format has them. Letting a lobby of strangers self-select roles reliably produces four of one and none of another.
Readiness checks
The gap between “accepted the invitation” and “can play at 7pm” is where multiplayer sessions die. Close it with an explicit checklist, run before the day, treated as all-or-nothing:
- Confidentiality terms accepted.
- Build installed. Not “has the key”: installed.
- Build launched successfully at least once on the machine they will use.
- Connection checked against the actual session route.
- Voice checked on the actual client the session will use.
Partial readiness is the dangerous state, because it looks fine on a dashboard. Four out of five checks with the build never launched is a seat that will fail at start time. Treat readiness as one boolean derived from the items, never as a percentage.
Run this at T–24 so there is still time to promote a standby, then re-check at T–2.
Session evidence: capture the conditions, not just the play
The recording is the least ambiguous part of your evidence and the least useful on its own. What makes a session analysable months later is the surrounding record:
- A timestamped log of events (technical failures, disconnects, replacements, build problems, conduct issues) with a severity on each.
- Neutral operator observations written during play, not reconstructed afterwards.
- The funnel: awarded, accepted, ready, attended, no-show, released, replaced. It tells you whether the session you analysed is the session you designed.
- Corrections as new entries. Never edit the timeline; append to it.
- Every protocol deviation, however small it seemed at the time.
Write observations as what happened, not what it meant. “Three players walked past the objective marker twice” survives review. “The objective marker is unclear” is already a finding, and it will be argued with.
Post-session interviews
The session tells you what happened. The interview tells you why, and it is the part most often cut when the session runs long. Protect the time in the schedule.
- Do it immediately. Recall degrades fast, and by the next day people report a tidied-up version of their own reasoning.
- Anchor to observed moments. “At 14 minutes you stopped and looked around for a while. What was going on there?” gets you a reason. “What did you think of the level?” gets you a review.
- Interview individually where it matters. A group debrief after a multiplayer session converges on whatever the most confident player said first.
- Make criticism costless and say so. If a participant thinks negative feedback affects future selection or payment, you will get politeness instead of data. The fix is structural: rewards that cannot be reduced, and telling people that.
- Ask what they expected. The difference between expectation and outcome is where most usable multiplayer findings live.
Before you write the findings
Attach a scope to every finding: one session, one cohort, one build, one region. Attach a confidence, and whether the evidence was consistent, mixed or contradicted. Then name the next test, and allow “replicate this before acting on it” to be the answer, because for a single session it usually is.
A finding that cannot be traced back to a specific observation, event or readiness record is an opinion that attended a playtest. It is fine to have opinions. Just do not ship them in the same document, under the same formatting, as the evidence.
Playtest Live