# V5 Capacity & Load Plan

plan-id: `v5-capacity-load.v1`

Status: Draft for review — fills audit gap V1_V7_PLAN_SET_AUDIT_2026-06-12 §6.2
("no concurrency/CCU targets or load-test plan for a 64-player persistent
world; GameLift+EKS named, never sized").

Owners: V5 Capacity Lead (SRE) — fleet sizing, autoscaling, cost model. V5
Online Services Lead — per-service SLOs. V5 Netcode Lead — game-server session
budgets and load-test pass criteria. Live-Ops Producer — CCU forecast sign-off.

Grounding: `V5/V5_DEPENDENCIES.md` §16 (AWS multi-region us-east / us-west /
eu-west / ap-northeast / ap-southeast + Alibaba CN; GameLift game servers; EKS;
MSK Kafka; ClickHouse on EKS), `V5/V5_ARCHITECTURE.md` §"Online Services
Backend" (16 services), `V5/V5_features.md` §"Online Services & Multiplayer
Infrastructure" (64-player Bureau HQ dedicated hub, 80 ms matchmaking ping
ceiling, 5%/20% packet-loss envelope, sub-2 s host migration, 500-event/2 KB
client telemetry buffering), launch date 2027-03-01 per
`V5/live-service/season-1-live-service-manifest.json`.

---

## 1. CCU Targets

### 1.1 Launch-peak derivation from open-world comparables

Planning assumption adopted 2026-06-12 — comparable Steam launch-peak
concurrents (public SteamDB figures, rounded):

| Title (class)                          | Steam launch peak CCU |
| -------------------------------------- | --------------------- |
| Cyberpunk 2077 (established-studio IP) | ~1,050,000            |
| Elden Ring (established IP)            | ~953,000              |
| Hogwarts Legacy (licensed mega-IP)     | ~880,000              |
| Starfield (established studio, new IP) | ~330,000              |
| Dying Light 2 (sequel, mid-tier)       | ~275,000              |
| New-IP AAA open-world median           | 100,000–300,000       |

V5 is a **new IP**. The mega-IP comparables are ceilings, not bases. Base case:
Steam peak of **60,000 CCU** (lower half of the new-IP band, reflecting no
franchise pull, offset upward by the five-cells-in-one-box pitch).

V5 ships on 9 platforms with full crossplay. Planning assumption adopted
2026-06-12: Steam ≈ 40% of total concurrents for a crossplay multiplatform
title (PC-heavy genre mix pulls Steam share up; console install base pulls it
down).

```
Global launch peak CCU = 60,000 / 0.40 = 150,000 CCU
Design point (1.5x surge headroom) = 225,000 CCU
Sustained launch-month average (peak:avg 2.5:1) = 60,000 CCU
Post-launch steady state (month 4+, 30% of peak) = 45,000 CCU
```

**Targets: provision for 225k CCU, alert-test at 150k, cost-model at 60k
average.**

### 1.2 Online concurrency mix

V5 is offline-capable and single-player heavy (features§"V5 Product Promise");
not all CCU touch multiplayer. Planning assumption adopted 2026-06-12, to be
recalibrated against closed-beta telemetry (see
`V5/docs/planning/beta-program.md`):

| Slice of CCU                                                        | Share | At 150k peak |
| ------------------------------------------------------------------- | ----- | ------------ |
| Offline / online-but-solo (services-light: auth, telemetry, sync)   | 65%   | 97,500       |
| Bureau HQ 64-player hub presence                                    | 14%   | 21,000       |
| Dedicated-server MP (Urban Freeroam, Frontier Posse, Capital Raid, sci-fi PvP, twitch PvP) | 16% | 24,000 |
| Listen-server co-op + LAN (no game-server cost)                     | 5%    | 7,500        |

Online-session CCU = 35% of total = 52,500 at peak / 78,750 at design point.

### 1.3 Regional split

Planning assumption adopted 2026-06-12 (comparable AAA crossplay titles):
us-east 30%, eu-west 30%, us-west 15%, ap-northeast 15%, ap-southeast 10%. CN
(Alibaba Cloud) is sized separately by the CN publishing partner and is out of
scope here.

---

## 2. GameLift Fleet Sizing — 64-Player Dedicated Hubs

### 2.1 Per-session budget

Planning assumption adopted 2026-06-12, **validated as a hard gate in technical
alpha** (beta-program.md §Technical Alpha): one UE5.5 dedicated hub server
session (64 players, 30 Hz server tick, Iris replication per ADR 0014) budgets:

- **4 vCPU** sustained (server frame p99 ≤ 33.3 ms at 64 players)
- **6 GB RSS** (hub world + 64 player states + workshop data-only mod sets)
- **Egress ≈ 16.4 Mbps/session** (64 players × 256 kbps replication budget)

### 2.2 Sessions per instance

Instance type: **c6i.12xlarge** (48 vCPU, 96 GB) — GameLift-supported, best
vCPU-$/perf in the c6i family at this size.

```
CPU bound:    floor((48 - 2 reserved for OS/agent) / 4) = 11 sessions
Memory bound: floor((96 - 8 reserved) / 6)              = 14 sessions
Network:      10 x 16.4 Mbps = 164 Mbps  << instance cap — not binding
Packing cap (scheduling + noisy-neighbor headroom):       10 sessions/instance
```

**10 hub sessions (640 player-slots) per c6i.12xlarge.**

Small-mode servers (16-player rooms: Capital Raid squads, twitch PvP lobbies,
private rooms from the server browser): 1 vCPU / 1.5 GB per session → packing
cap **40 sessions (640 slots) per c6i.12xlarge** (same slot density; memory
bound 58, CPU bound 46, capped at 40 for headroom).

### 2.3 Fleet math at design point (225k CCU)

```
Hub players      = 225,000 x 14% = 31,500 → 31,500/64 = 493 sessions
Hub instances    = ceil(493 / 10) = 50  → +20% buffer = 60 instances global
Dedicated-MP     = 225,000 x 16% = 36,000 → 36,000/16 = 2,250 sessions
MP instances     = ceil(2,250 / 40) = 57 → +20% buffer = 69 instances global
Total game-server fleet at design point ≈ 129 c6i.12xlarge, split by §1.3:
  us-east 39, eu-west 39, us-west 19, ap-northeast 19, ap-southeast 13
```

Buffer policy: **20% available-session buffer** at all times (GameLift
target-tracking on `PercentAvailableGameSessions = 20`). FleetIQ Spot/On-Demand
mix **60/40** — Spot interruption is acceptable for small-mode rooms (players
re-matchmake), On-Demand pinned for hub sessions (64-player presence must not
evaporate; hub sessions are placed only on the On-Demand portion).

First-party relays (PSN, XBL, Steam Datagram Relay) front platform-pool
sessions per features§"Online Services" degradation rules; the GameLift fleet
above carries all cross-platform sessions and is sized as if it carried 100% —
relay offload is upside, not a dependency.

---

## 3. Per-Service Capacity SLOs (16 services)

All rates at the **225k design point**. Availability SLOs are monthly.
Latency = p99 server-side. Derivations noted inline; every service is
horizontally scaled on EKS (HPA, see §4).

| Service                | Peak load (derivation)                                                                 | p99     | Availability | Degradation behavior (per spec)                              |
| ---------------------- | -------------------------------------------------------------------------------------- | ------- | ------------ | ------------------------------------------------------------ |
| auth                   | 250 RPS login (patch-storm: 50% of CCU reconnect in 10 min = 112,500/600 = 188/s, rounded up) + 250 RPS refresh (15-min JWT: 225k/900 s) | 300 ms  | 99.95%       | Outage drops players to offline-deterministic play           |
| friends                | 3,750 RPS presence (1 update/min/online user = 225k/60)                                 | 400 ms  | 99.9%        | Stale presence acceptable; no play impact                    |
| parties                | 400 RPS (≈5% of online users in party ops per minute)                                   | 400 ms  | 99.9%        | Party features unavailable; solo queue still works           |
| matchmaking            | 150 tickets/s (re-queue churn: 2,250 sessions / 20-min avg = 1.9 sessions/s x 16 players = 30/s steady; 5x surge factor) ; TTM p95 < 60 s | 500 ms | 99.9% | Clear retry state, never a hang (features§"Online Services") |
| leaderboards           | 30 writes/s (match completions incl. anti-cheat re-validation) + 500 reads/s            | 250 ms  | 99.5%        | Score posts queue; reads serve cached                        |
| replays                | 5 uploads/s x 20 MB (≈10% of completed matches) → S3; 50 RPS metadata                   | 1 s     | 99.5%        | Upload retries client-side                                   |
| telemetry              | 112,500 events/s (225k x 0.5 events/s/client); 7,500 batch-POSTs/s (30 s batching)      | 800 ms  | 99.9% ingest | Client buffers 500 events / 2 KB per batch, flushes on recovery |
| anti-cheat             | 2,600 signals/s (78,750 online x 2 signals/min); ML classifier consumes Kafka stream    | async   | 99.5%        | Fail-open to post-match re-validation; never blocks play     |
| crash-reporting        | 10 dumps/s sustained, 100/s after a bad patch (sizing case)                             | 2 s     | 99.0%        | Client retries with backoff                                  |
| companion-app-bridge   | 2,250 RPS (≈5% of CCU concurrently in app x 0.2 RPS)                                    | 400 ms  | 99.5%        | Read-only; app shows stale data                              |
| mindpalace-cloud-sync  | 500 RPS (sync on device/session transition)                                             | 600 ms  | 99.9%        | Blocks cross-device pickup, not play                         |
| workshop               | 1,500 RPS catalog browse (downloads offloaded to CloudFront); 5 publishes/s             | 500 ms  | 99.5%        | Subscribed content already local; browse-only outage         |
| compliance-dsar        | Batch; 100 requests/day                                                                 | n/a     | 99.0%        | SLA measured in days (30-day GDPR clock), not ms             |
| live-service-calendar  | 2,000 RPS (edge-cached ≥95% hit ratio; origin ≤100 RPS)                                 | 200 ms  | 99.9%        | Clients cache last-known calendar                            |
| balance-ledger         | 300 RPS public reads (cached); <1 write/s                                               | 300 ms  | 99.5%        | Read-only public surface                                     |
| faction-rep            | 200 RPS (Sci-Fi cell writes on mission resolution)                                      | 400 ms  | 99.5%        | Buffered in cell save; reconciles on recovery                |

Error-budget policy: a service that burns >50% of its monthly budget in 7 days
freezes feature deploys for that service until burn returns under curve
(Live-Ops SRE Lead owns enforcement).

---

## 4. Autoscaling Policies

**GameLift fleets** (per region):

- Target-tracking: `PercentAvailableGameSessions = 20%`.
- Scale-out: available sessions < 15% for 2 consecutive minutes → add
  instances (step: +10% of current fleet, min 2 instances). Instance warm-up
  budget 8 minutes (AMI bake includes server binary; no download at boot).
- Scale-in: available sessions > 30% for 20 consecutive minutes → remove
  instances with **session protection on** (instance drains; never kills an
  occupied hub).
- Floor: launch week floor = 50% of design-point fleet held warm regardless of
  load (absorbs the 00:00 UTC 2027-03-01 spike faster than scale-out can).
- Spot interruption: small-mode rooms only; hub sessions On-Demand-pinned (§2.3).

**EKS services**: HPA on CPU 60% target plus per-service custom metric (RPS per
pod at 70% of load-tested ceiling). Min replicas = 3 per region (zonal spread).
Cluster Autoscaler headroom: 15% spare node capacity. Telemetry edge service
additionally scales on Kafka producer-buffer depth (>10 s of lag → scale out).

**ClickHouse / Kafka**: not autoscaled. Provisioned to design point (§5),
capacity-reviewed monthly by the V5 Capacity Lead; expansion is a planned
change, not an autoscaling event.

---

## 5. Load-Test Plan

**Tooling**: k6 (HTTP service load, distributed runners on a dedicated EKS
load cluster); custom UE Gauntlet headless bot clients ("GauntletLoad") for
game-server fill — bots run real client netcode and replication, scripted
movement/combat loops per cell; `tc netem` on bot hosts for latency/loss
injection; GameLift FlexMatch latency simulation for matchmaking-policy tests.
All runs against a production-shaped staging environment in us-east + eu-west
(2-region minimum to exercise cross-region paths).

Scenarios — each runs green **twice** before the open-beta stress weekend and
once on the launch-candidate build:

| ID    | Scenario                                                                                                                       | Pass criteria (tied to specced envelopes)                                                                                                          |
| ----- | ------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| LT-1  | **Login storm** — 250 RPS login + 250 RPS refresh for 30 min (patch-day reconnect, §3)                                          | auth p99 ≤ 300 ms; error rate < 0.1%; zero cascading failure into friends/parties                                                                    |
| LT-2  | **Hub fill** — 500 hub sessions × 64 bots (32,000 bots), 2-hour soak                                                            | Server frame p99 ≤ 33.3 ms per session; zero session crashes; RSS ≤ 6 GB/session; egress ≤ 16.4 Mbps/session                                          |
| LT-3  | **Matchmaking surge** — 150 tickets/s with simulated regional ping spread up to 80 ms                                            | TTM p95 < 60 s; no match formed above the 80 ms ping ceiling (features§"Online Services"); retry state (never hang) when queue saturated             |
| LT-4  | **Host-migration storm** — 1,000 concurrent listen-server co-op sessions; kill 10% of hosts within a 60 s window, repeat 5 waves | Migration success ≥ 99.5%; migration time p95 < 2 s (features§"Network Quality, LAN, and Server Browser"); no save corruption in any migrated session |
| LT-5  | **20%-loss degradation** — netem 20% loss + 80 ms RTT on 25% of bots in twitch-PvP and open-world MP sessions                     | Twitch PvP stays within 8-frame rollback window without desync; open-world MP hit validation stays lag-compensated; session remains playable at 5% loss and degrades gracefully (no disconnect cascade) at 20% per arch§"Observability, Evaluation, Release Gates" |
| LT-6  | **Telemetry flood** — 112,500 events/s for 1 hour, then kill ingestion for 10 min, then recover                                  | Zero event loss end-to-end (client 500-event/2 KB buffers flush on recovery); Kafka consumer lag drains < 15 min after recovery                       |
| LT-7  | **Replay surge** — 20 uploads/s × 20 MB for 15 min (4× nominal)                                                                  | Upload success ≥ 99%; no impact on matchmaking/leaderboard SLOs (bulkhead check)                                                                      |
| LT-8  | **Full-region failover** — drain us-east at 60% design load                                                                       | Surviving regions absorb load within autoscaling SLAs; matchmaking re-homes within 5 min; error budget burn < 10% of monthly during the drill          |

A failed scenario blocks launch the same way a Launch Readiness gate does
(features§"Launch Readiness"); LT results attach to the launch-readiness
evidence bundle.

---

## 6. Infrastructure Cost Model (skeleton)

All prices are us-east-1 list-price snapshots, planning assumption adopted
2026-06-12; the model's purpose is structure and order-of-magnitude, refreshed
quarterly by the V5 Capacity Lead with Finance.

**Game servers (GameLift)**

```
c6i.12xlarge on-demand ≈ $2.04/hr; Spot ≈ 30% of on-demand ≈ $0.61/hr
Blended (60% Spot / 40% OD)        = 0.6 x 0.61 + 0.4 x 2.04 = $1.18/hr
Player-slots per instance          = 640 (both hub and small-mode packing, §2.2)
Cost per occupied player-hour      = 1.18 / (640 x 0.70 occupancy x 0.83 buffer factor)
                                   = 1.18 / 372 ≈ $0.0032
Share of CCU on dedicated servers  = 14% hub + 16% MP = 30%
Game-server cost per total CCU-hr  = 0.30 x 0.0032 ≈ $0.00096 ≈ $0.0010
```

**ClickHouse on EKS (telemetry warehouse)**

```
Avg ingest at 60k CCU = 60,000 x 0.5 events/s x 350 B = 10.5 MB/s
Raw/day = 10.5 x 86,400 ≈ 907 GB; compressed (10:1) ≈ 91 GB/day
90-day hot retention ≈ 8.2 TB → provision 12 TB gp3
Compute: 6 x m6i.4xlarge x $0.768/hr = $4.61/hr ≈ $3,365/mo
Storage: 12,288 GB x $0.08/GB-mo ≈ $983/mo
Total ≈ $4,350/mo → / (60,000 CCU x 730 hr) ≈ $0.0001 per CCU-hr
```

**Kafka (MSK)**

```
Peak ingress 39.4 MB/s (design point) x 3x replication ≈ 118 MB/s cluster write
6 x kafka.m5.2xlarge x $0.84/hr = $5.04/hr ≈ $3,680/mo + storage ≈ $4,500/mo
≈ $0.0001 per CCU-hr
```

**Databases + cache (16 services)**

```
Avg per service (RDS multi-AZ r6g mix + share of ElastiCache) ≈ $790/mo
16 x $790 ≈ $12,600/mo ≈ $0.0003 per CCU-hr
```

**CDN + S3 (replays, workshop, patches — excl. launch-day patch spike)**

```
≈200 TB/mo egress x $0.06/GB committed ≈ $12,000/mo + S3 ≈ $2,000/mo
≈ $0.0003 per CCU-hr
```

**EKS service compute + observability stack** ≈ $9,000/mo ≈ $0.0002 per CCU-hr

**Multi-region adder** (cross-region replication, NAT, inter-AZ, 5 regions
operating floors): **+15%** on the subtotal.

```
Subtotal per CCU-hr  = 0.0010 + 0.0001 + 0.0001 + 0.0003 + 0.0003 + 0.0002
                     = $0.0020
With multi-region    = 0.0020 x 1.15 ≈ $0.0023 per CCU-hour
Launch-month (60k avg CCU): 60,000 x 730 x 0.0023 ≈ $101,000/mo
Steady state (45k avg CCU): ≈ $76,000/mo
Launch-week design-point burst adds ≈ $25k one-time (fleet floor at 50%, §4)
```

Out of scope for this skeleton (tracked by Finance with the Live-Ops Producer):
CN region (partner-borne), voice infrastructure (provider-dependent, see
`V5_DEPENDENCIES.md`), bandwidth for the day-one patch, and people costs.

---

## 7. Review cadence

- Recalibrate §1 mix and §2 session budgets after technical alpha (2026-10),
  again after closed beta (2026-12) — owner: V5 Capacity Lead.
- LT-1..8 green twice before the open-beta stress weekend (2027-01-28) and once
  on the launch candidate — owner: V5 Netcode Lead.
- Cost model refresh quarterly and at each beta gate — owner: V5 Capacity Lead
  + Finance.
