plan-id: v5-beta-program.v1
Status: Draft for review — fills audit gap V1_V7_PLAN_SET_AUDIT_2026-06-12 §6.2 ("no beta or technical alpha before a straight-to-launch netcode envelope claim"). The launch-readiness gates (features§"Launch Readiness") currently assert validated envelopes — 80 ms matchmaking ceiling, 5%/20% packet-loss tolerance, sub-2 s host migration — with no plan for validating them against real players on real networks. This program is that plan.
Launch: 2027-03-01 (V5/live-service/season-1-live-service-manifest.json).
All cohort sizes and dates are planning assumptions adopted 2026-06-12; dates
are owned by the Live-Ops Producer, criteria by the named gate owners.
Phase map:
| Phase | Window | Cohort | Primary de-risk |
|---|---|---|---|
| Technical alpha | 2026-09-17 → 2026-10-04 | 5,000 invited (NDA) | Netcode envelope + hub server budgets, with real players |
| Closed beta | 2026-11-12 → 2026-12-06 | 75,000 invited | Retention-shaped telemetry, economy, workshop, matchmaking quality |
| Open beta stress weekend | 2027-01-28 → 2027-01-31 | Uncapped, target 200k+ peak | Capacity plan at design point; launch-day ops rehearsal |
Buffer math: open beta ends 2027-01-31, leaving 4 weeks to launch — enough for one client hotfix cycle through expedited platform cert (5–10 business days typical) plus a fleet-config iteration, and deliberately not enough for feature work: open beta is validation, not development. Anything open beta finds that needs >2 weeks of engineering triggers the launch-slip conversation explicitly, with the Live-Ops Producer owning the call — that is the honest function of a beta this close to ship.
Phase 1 — Technical Alpha (2026-09-17 → 2026-10-04)#
Three weekends, PC-only (Steam build under NDA; consoles excluded to avoid early cert — platform-cert risk is carried by closed beta instead). Content: Bureau HQ hub, one dedicated-MP mode per cell group (Urban Freeroam, twitch PvP), listen-server co-op missions, offline single-player slice for the reconcile path. Gate owner: V5 Netcode Lead.
Cohort: 5,000 — sized for statistics, not scale: at 5,000 players with a measured regional/ISP spread (recruited deliberately across the 5 AWS regions' catchments, including high-latency and lossy-network segments — target ≥15% of cohort on connections worse than 40 ms / 1% loss), every envelope metric below gets ≥10^4 session-samples per weekend, enough to resolve p95s with tight confidence intervals. Scale is not the goal; LT-1..8 bot tests (capacity plan §5) own scale.
Entry criteria
- LT-2 (hub fill) and LT-4 (host-migration storm) green against staging with bots — real players never debug what bots can catch first.
- Crash-free session rate ≥ 97% on the alpha build in internal playtests.
- Telemetry pipeline capturing the full netcode metric set (per-session RTT, jitter, loss, migration timings, server frame times).
Success / exit criteria (each maps to a claimed envelope)
| Claimed envelope (spec) | Alpha validation bar |
|---|---|
| Playable at 5% loss; graceful at 20% (arch release gates) | Sessions with measured 4–6% loss: quality-grade HUD correct, disconnect rate < 2x clean-network baseline; 15–25% loss: degradation states engage, no crash, no save loss, disconnect is clean |
| 80 ms matchmaking ping ceiling | 0 matches formed above ceiling; players at 70–80 ms report match quality ≥ 4/5 in the in-client survey (n ≥ 300) |
| Sub-2 s host migration | Real-network migration p95 < 2 s, success ≥ 99.5% across ≥ 2,000 organic + prompted host drops |
| 64-player hub server budgets (capacity plan §2.1) | Server frame p99 ≤ 33.3 ms and RSS ≤ 6 GB on full real-player hubs (≥ 50 full-hub hours); if breached, the capacity plan's fleet math reflows before closed beta — this is the gate that turns the 4 vCPU/6 GB planning assumption into a measurement |
| 8-frame rollback (twitch PvP) | Desync rate < 0.1% of rounds; one-sided-advantage complaints < 1% of surveyed matches |
| Offline reconcile | 100% of offline-session reconciles either apply cleanly or fail-loud to the quarantine path; 0 silent wallet/progress divergence (economy doc §4) |
What it de-risks: the netcode envelope claims stop being simulation-only; the GameLift session budget becomes empirical; host migration is proven on consumer NAT/Wi-Fi reality rather than netem.
Exit decision 2026-10-09: V5 Netcode Lead signs each row with data attached, or the row's remediation lands on the closed-beta entry list.
Phase 2 — Closed Beta (2026-11-12 → 2026-12-06)#
Three-and-a-half weeks continuous (not weekend-gated — economy and retention need continuous time), PC + 2 console platforms (first cert exposure on a beta build). Content: all five cells' opening hours, hub, full dedicated-MP mode set, workshop with free publishing live (marketplace in supervised pilot with the six Year-1 listings' creators per the manifest), persistent-economy surfaces. Gate owners: Live-Ops Producer (overall), Economy Designer (economy gates), Trust & Safety Lead (workshop gates).
Cohort: 75,000 invited (waves of 25k), split by the capacity plan's regional mix (us-east 30%, eu-west 30%, us-west 15%, ap-northeast 15%, ap-southeast 10%) so per-region matchmaking pools are realistic.
Entry criteria
- All technical-alpha exit rows signed or remediated.
- Economy simulations EC-S1..S6 green (economy doc §6).
- Workshop moderation pipeline staffed end-to-end (ML pre-screen + report queues + human moderators + DSA appeal path) at beta scale.
- DMCA agent registered (compliance doc §7 gate 1).
- Incident-response rotations live in beta-severity form; status page live.
Telemetry gates / exit criteria
| Domain | Gate |
|---|---|
| Stability | Crash-free session rate ≥ 99.0% on PC, ≥ 98.5% on consoles (launch bar is 99.5% — beta bar is on-glidepath, not at-launch) |
| Matchmaking | TTM p95 < 90 s in every region at beta population (launch bar 60 s at launch population); match-quality survey ≥ 4/5 median |
| Services | No service below its availability SLO floor for the final 2 weeks of beta |
| Economy | AAMW baseline curves captured per cohort (this beta creates the §3 baselines); live sink/faucet ratio within 0.85–1.05; 0 unexplained ledger-invariant breaches in the final 10 days |
| Workshop | Moderation queue time p90 < 24 h; ML pre-screen false-positive rate < 5% measured against human review; 0 data-only sandbox violations |
| Hub | 64-player hubs sustain real-population churn (join/leave storms at hour boundaries) within server budgets |
| Retention proxy | D7 return rate of wave-1 invitees ≥ 35% (planning assumption benchmark for invited-beta populations; below 30% triggers a product review, not just an ops review) |
What it de-risks: matchmaking quality at population scale, the economy's baseline curves and tuning, moderation throughput, console cert surprises caught 3+ months early, the §1.2 capacity-plan online-mix assumption (the 35% online share gets its first real measurement here and the capacity plan recalibrates).
Phase 3 — Open Beta / Server-Stress Weekend (2027-01-28 → 2027-01-31)#
Open to everyone, all 9 platforms if cert timing allows (platforms whose beta cert misses the window are excluded rather than slipping the weekend — the stress goal is population, not platform completeness). Content: a bounded slice (hub + 2 MP modes + one cell's opening) — small enough to cut from the launch-day patch, big enough to load every service. Gate owners: V5 Capacity Lead (load), Live-Ops SRE Lead (ops rehearsal).
Target: 200,000+ peak CCU — i.e., at or above the 150k launch-peak forecast and approaching the 225k design point. Marketing owns demand generation; if organic demand under-shoots 150k, the bot fleet (GauntletLoad) tops up the difference so the infrastructure number is hit either way — with the honest caveat recorded in the exit report that bot-topped load validates infrastructure but not human behavior patterns.
Entry criteria
- LT-1..8 green twice (capacity plan §5).
- All closed-beta exit gates signed or explicitly risk-accepted by the Live-Ops Producer in writing.
- Incident drill windows 1–3 complete (incident doc — the stress weekend runs on the drilled rotations; drill window 4 follows it).
- Day-one patch infrastructure (CDN paths, launcher flows) is what serves the open-beta client — the download surge is one of the tests.
Exit criteria
- Autoscaling holds the 20% available-session buffer through the Friday-night and Saturday-peak ramps with no manual fleet intervention beyond the pre-planned floor changes.
- Every per-service SLO in the capacity plan §3 table holds at observed peak; any breach gets a sized remediation that fits the 4-week window, or it goes to the launch-slip conversation.
- One deliberate game-day executed live during the weekend: a scheduled region drain (us-west, lowest-traffic window) exercising the region-outage runbook against real traffic. Pass = matchmaking re-homes within 5 min, status-page comms hit their SLAs.
- Zero SEV-1s caused by load (SEV-1s caused by the deliberate game-day are scripted and excluded); all SEV-2s mitigated within the 4 h matrix target.
- Cost telemetry from the weekend reconciles the §6 cost model within ±30% (a bigger error means the launch-month budget is wrong — Finance review).
What it de-risks: the launch-day shape itself — surge, download storm, autoscaling, on-call execution, status comms — five weeks before it happens, while there is still one full fix-and-cert cycle of runway.
Cross-phase rules#
- Every phase ships the real telemetry, anti-cheat, and economy pipelines — no beta-only stubs; the pipelines are themselves under test (zero-stub policy applies to test infrastructure too).
- Beta progression wipes before launch, stated in every beta ToS up front; cosmetic "beta veteran" reward at launch (cheap goodwill, costs nothing).
- Each phase ends with a written exit report (gate table with measurements, signed by owners) attached to the launch-readiness evidence bundle — the Launch Readiness gates in features§"Launch Readiness" cite these reports as the "validated" evidence they currently assert without a source.