Oshun Platform · Planning

V1 Capacity Model — Living Scenes, Psyche, BFF

All rows are _planning assumptions adopted 2026-06-12_ with derivations; they define the design point the fleet is sized for (≈ GA+90 days), plus the smaller GA-week point used for initial procurement.

8sections10 minread9tables

On this page

Status: planning baseline, created 2026-06-12 to close V1_V7_PLAN_SET_AUDIT_2026-06-12.md §6.2 (V1: "no capacity/load numbers for GPU-bound Living Scenes"). All spec budgets cited below are the product's own committed numbers; everything labeled planning assumption adopted 2026-06-12 is replaced by measurement during beta (§8).

1. Spec inputs this model is derived from#

Budget Value Source
Living Scenes first-frame latency p50 ≤ 400 ms, p95 ≤ 800 ms V1/features.md:4411
Segment generation latency (per 4–12 s segment) p50 ≤ 6 s, p95 ≤ 12 s V1/features.md:4412-4413
Pre-warm lookahead ≥ 2 segments V1/features.md:3723-3727
Composition tier caps free ≤ 6 segments / ≤ 90 s; paid ≤ 24 segments / ≤ 8 min V1/features.md:3803-3808
Psyche end-to-end voice p50 ≤ 500 ms, p95 ≤ 900 ms, p99 ≤ 1500 ms V1/features.md:2220-2222
Psyche first-token p50 ≤ 350 ms, p95 ≤ 700 ms V1/features.md:2223
Avatar-frame latency p50 ≤ 80 ms V1/features.md:2224
ASR final-text latency p50 ≤ 150 ms after end-of-speech V1/features.md:2225-2226
Voice cue effect p95 ≤ 800 ms; text cue p95 ≤ 300 ms V1/features.md:3697-3700
GPU cloud RunPod (A100/H100/RTX 4090) V1/DEPENDENCIES.md:255
Segment generators Hunyuan/WAN-class with per-segment failover V1/features.md:3740-3742
Backpressure ladder extend segment → reduce resolution → lengthen blend → still+narration V1/features.md:4420-4422
Nyx base-render caching once per location-bucket per night V1/features.md:4072-4074
Shared MP4 cache 30 days post-share, then re-render on demand V1/features.md:4196-4199

2. Launch concurrency assumptions (explicit)#

All rows are planning assumptions adopted 2026-06-12 with derivations; they define the design point the fleet is sized for (≈ GA+90 days), plus the smaller GA-week point used for initial procurement.

Quantity GA-week Design point Derivation
Registered users 25k 130k beta exit ≈ 10k invitees (LAUNCH_TIMELINE Phase 2) + canary ramp; design point assumes 100k MAU at GA+90
DAU 8k 20k DAU/MAU 20 % — mid-range for a daily contemplative-practice product (habit loop pushes above social-app 15 %, below messaging 50 %)
Peak concurrent users (PCU) 640 1,600 → sized at 2,000 peak-hour concurrency 8 % of DAU (single dominant evening peak for a contemplative product); sized at 2,000 for headroom
Concurrent Living Scenes sessions 64 200 10 % of PCU in a scene at peak — Living Scenes is the flagship surface, but most peak users are in text/browse surfaces
Free : paid scene mix 75 : 25 75 : 25 standard freemium attach at launch; re-measured weekly
Concurrent Psyche voice sessions 96 300 15 % of PCU using voice mode (voice is high-friction in public settings)
Concurrent avatar sessions 24 75 25 % of voice sessions escalate to avatar
BFF sustained request rate 192 RPS 600 RPS 0.3 req/s per concurrent user (navigation, polling, autosave, telemetry batch)
BFF launch-spike multiplier (3,000 RPS) launch-announcement traffic shape; verified by LT-2
Concurrent WebSocket connections 768 2,400 PCU × 1.2 (multi-tab/multi-device)

3. Living Scenes GPU fleet sizing#

3.1 SKU selection — the spec budget forces it#

The segment budget (p50 ≤ 6 s for a 4–12 s segment) requires a distilled few-step Hunyuan/WAN-class video model at 480–720p. Relative throughput (planning assumption adopted 2026-06-12; FP8/TensorRT-class inference, H100 ≈ 2–2.5× A100 on video-diffusion workloads):

SKU Expected p50 per segment Verdict
H100 80GB ≈ 6 s Only SKU that meets the live-tier budget. Mandated for live rendering.
A100 80GB ≈ 12–15 s Breaches p50 budget; eligible only for offline work: re-bakes, reduced-motion stills, watermark bake, thumbnails
RTX 4090 VRAM-constrained for the model class Dev/test only

This also serves the determinism release gate (V1/features.md:4193): one SKU plus one driver/CUDA/container digest per workflow class (RISK_REGISTER R-08).

3.2 Renders-per-GPU-hour math (from the spec's own budget)#

  • Segment renders/GPU-hour at p50 budget: 3600 s ÷ 6 s = 600 renders/GPU-hr.
  • Sustainable effective throughput at 75 % target utilization (scheduling gaps, model load, p95 tail): 600 × 0.75 = 450 renders/GPU-hr.
  • Mean segment duration: 9 s (mid-band of the 4–12 s spec range; planning assumption adopted 2026-06-12).
  • Per-session GPU occupancy = gen time ÷ playback time = 6 ÷ 9 = 0.67 GPU steady-state at p50 (p95 worst case 12 ÷ 9 = 1.33 — absorbed by the ≥ 2-segment pre-warm lookahead, not by fleet overprovisioning).

3.3 Cache and replay offset#

30 % of session-seconds need no live generation (planning assumption adopted 2026-06-12). Two sources: Nyx briefings materialize once per location-bucket (V1/features.md:4072-4074) — ≈ 20 % of sessions × 80 % cache hit — and replays of kept artifacts are served from the 30-day MP4 cache (V1/features.md:4196-4199) — ≈ 15 % of sessions fully cached.

3.4 Fleet formula and numbers#

text
fleet = sessions × occupancy × (1 − cache_offset) ÷ utilization + warm_pool
Point Sessions Compute Live fleet Warm pool Total H100
GA-week 64 64 × 0.67 × 0.70 ÷ 0.75 = 40 40 8 48 → procure 50
Design point 200 200 × 0.67 × 0.70 ÷ 0.75 = 125 125 10 135 → procure 130 + burst

Warm pool / first-frame budget: The p50 ≤ 400 ms first-frame budget cannot be met by cold generation. Mechanism: (a) the opening segment of every template variant is pre-rendered and cached (Tara breath-cycle anchors, Nyx celestial anchors, etc., V1/features.md:3797-3802), so first frame is a cache read + stream start; (b) a warm idle-worker pool of max(8, 5 % of fleet) absorbs session-start bursts so the second (first personalized) segment begins generation immediately.

Procurement split (feeds RISK_REGISTER R-01):

  • 60 % RunPod Secure Cloud committed (78 of 130 at design point),
  • 40 % RunPod on-demand,
  • plus 20 % of peak (26 H100) warm at AWS EC2 P5 as the cross-provider failover reserve, same image digest, registered as drained in the Conductor worker pool.

3.5 Scaling policy#

Diurnal: scale-in floor = 30 % of peak fleet (model-load latency makes cold-adds slow); scale-out signal = queue depth (§6) and projected wait; predictive pre-scale 30 min ahead of the historical evening peak per region.

4. Psyche voice/avatar capacity#

4.1 Per-node session capacity#

Psyche runtime node = c7i.2xlarge-class (8 vCPU / 16 GB) — planning assumption adopted 2026-06-12. Per active voice session the node carries Opus encode/decode, VAD, event fan-out, and transcript sync ≈ 25 ms CPU per second of audio ≈ 2.5 % vCPU (planning assumption; derived from public Opus/WebRTC SFU benchmarks at 48 kHz mono). At a 60 % CPU target:

  • Voice sessions/node: 8 vCPU × 60 % ÷ 0.025 ≈ 190 → plan 150/node (jitter + reconnect-storm headroom).
  • Avatar sessions/node: viseme + expression streams roughly double per-session event throughput (V1/features.md:2202-2215) → 75/node.

4.2 Node and vendor sizing at design point#

Quantity Value Derivation
Concurrent voice sessions 300 §2
Psyche nodes 3 minimum (AZ-spread), surge to 6 300 ÷ 150 = 2 nodes of load; N+1 across 3 AZs
ElevenLabs contracted TTS concurrency 450 streams 300 × 1.5 safety factor; default ElevenLabs plan concurrency caps are far below this — an enterprise concurrency contract is a launch dependency (RISK_REGISTER R-02)
ASR concurrency (primary + fallback) 450 streams mirrors TTS; self-hosted Whisper-class fallback at 2 warm GPUs per R-02

4.3 Budget decomposition constraint#

The voice budget only closes with a fully streamed pipeline: ASR final ≤ 150 ms (V1/features.md:2225) + LLM first token ≤ 350 ms (:2223) already sums to 500 ms — equal to the entire p50 end-to-end budget (:2220). Therefore: ASR finalization must overlap LLM prefill, and TTS synthesis must start on the first clause with a flash-tier voice model (first-chunk ≈ 75–100 ms, planning assumption adopted 2026-06-12). Any provider or config that serializes these stages breaches p50 by construction — load tests assert overlap (§7 LT-1).

5. Unit economics (feeds RISK_REGISTER R-04)#

GPU price: H100 80GB ≈ $2.99/GPU-hr RunPod Secure Cloud (planning assumption adopted 2026-06-12; mid-point of recent published secure-cloud rates; AWS P5 failover capacity runs 2–3× this, which is why it is reserve, not primary).

Item GPU-seconds Cost
Free scene (6 segments × 6 s gen) 36 $0.03
Paid scene (24 segments × 6 s gen) 144 $0.12
Design-point fleet, peak hour (130 GPUs) ≈ $389/hr
Design-point monthly (diurnal avg 45 % of peak) 58.5 × $2.99 × 720 h ≈ $126k/mo
GA-week monthly (50 GPUs, same shape) $48k/mo

Kill criterion and per-user daily budgets are defined in RISK_REGISTER R-04.

6. BFF capacity and queue-depth thresholds#

6.1 BFF (Fastify, apps/oshun/bff, single BFF per V1/ARCHITECTURE.md:292)#

Quantity Value Derivation
Sustained target 600 RPS §2
Spike target 3,000 RPS (5×) §2
Pod design capacity 800 RPS/pod @ 2 vCPU, p99 ≤ 120 ms planning assumption adopted 2026-06-12 — conservative for Fastify JSON proxy + JWT verify + rate-limit lookup; validated, not trusted: LT-1 measures it
Pod count min 4 (N+2 over 3 AZs), HPA max 16 4 × 800 covers spike start while HPA reacts; 16 × 800 = 12,800 RPS ceiling ≈ 4× spike target
WebSocket connections 2,400 concurrent §2; held by Psyche gateway pods, not BFF worker pool
DB connections per-domain PgBouncer pools (V1/ARCHITECTURE.md:1210-1211), 20 server conns per domain pool planning assumption; prevents BFF HPA from exhausting Postgres

Per-client rate limiting at the BFF is already spec'd (V1/ARCHITECTURE.md:1069-1071); overload shedding order in RISK_REGISTER R-07.

6.2 Segment-render queue thresholds#

Let C = current live fleet size (one in-flight segment per GPU). Projected wait = depth × p50 gen time ÷ C.

State Condition Action
GREEN depth ≤ C normal
AMBER depth > 2C sustained 60 s, or projected wait > 6 s engage the spec backpressure ladder in order (extend current segment → reduce next-segment resolution within tolerance → lengthen blend window, V1/features.md:4420-4422); fire scale-out
RED depth > 4C, or projected wait > 9 s (one mean segment — the stall horizon given ≥ 2-segment lookahead) admission control: new free sessions are offered the reduced-motion variant or a queued start; paid sessions admitted until depth > 6C; in-flight sessions degrade per ladder, never killed
BLACK depth > 6C or fleet < 50 % of target all new sessions queued with honest wait estimate; still+narration fallback for in-flight (V1/features.md:4422); page SRE

Crisis frames are never shed or queued in any state — the crisis path (V1/features.md:3701-3704) has reserved capacity (2 GPUs held out of the admission pool).

6.3 Voice admission threshold#

If ElevenLabs concurrency utilization > 85 % of contract, new sessions start in text mode with a voice-upgrade offer (matches the spec's recovery behavior, V1/features.md:2245-2246) rather than risking mid-session budget breaches.

7. Load-test plan#

Tooling: k6 (HTTP + WebSocket scenarios) for BFF and Psyche envelope; Conductor replay harness driving the per-template fixture Scores (≥ 30 per template, V1/features.md:4466-4469) at controlled arrival rates against the real GPU fleet; Prometheus/Grafana (observability profile, V1/DEPENDENCIES.md §21) for measurement; fault injection at the provider adapters for chaos legs. Test types map to the named V1 quality matrix — load, spike, soak, memory-leak, performance, queue-throughput, failure, recovery, chaos, disaster-recovery (V1/features.md:6171-6173).

ID Type Scenario Pass criteria
LT-1 load / performance 1× design point (2,000 PCU, 200 scenes, 300 voice) for 2 h every §1 budget met under load: voice p95 ≤ 900 ms, first-frame p95 ≤ 800 ms, segment p95 ≤ 12 s, BFF p99 ≤ 800 ms read; error rate < 0.5 %; ASR/LLM/TTS overlap confirmed (§4.3)
LT-2 spike BFF 600 → 3,000 RPS in 60 s, hold 15 min 5xx < 2 % during ramp, < 0.5 % after; HPA stabilizes ≤ 5 min; no auth/entitlement path shed
LT-3 soak / memory-leak 0.7× design point for 8 h RSS growth < 10 % on BFF, Psyche, Conductor; no queue-depth creep; zero session-state leaks across reconnects
LT-4 queue-throughput segment arrivals at 1.5× render capacity for 30 min backpressure ladder engages in spec order; zero failed sessions; AMBER→GREEN recovery ≤ 10 min after load drops; admission control verifiably distinguishes free/paid
LT-5 chaos / failure-recovery kill 30 % of RunPod workers mid-load per-segment failover transparent (V1/features.md:3740-3742); AWS warm pool absorbs; segment p95 ≤ 1.5× budget during transition, recovered ≤ 15 min; failovers recorded in Render Envelope audit
LT-6 chaos (vendor) inject 50 % ElevenLabs failure + full ASR outage legs degradation ladder order honored with user disclosure (RISK_REGISTER R-02); text fallback within configured threshold; automatic re-upgrade observed
LT-7 disaster-recovery / restore per SLO_AND_DR.md §6 drill schedule RPO/RTO targets met

Gate linkage: LT-1 through LT-6 passing is an entry criterion for private beta wave 3 and for the canary ramp (LAUNCH_TIMELINE Phases 2–3). LT-7 is a GA go/no-go input.

8. Re-forecast cadence#

  • Weekly during beta: observed attach rates, session lengths, cache-hit rates, and measured renders/GPU-hr replace the §2/§3 assumptions; fleet plan re-derived from the same formulas.
  • At every canary promotion: capacity review signed by SRE Lead (the promotion criteria in LAUNCH_TIMELINE require headroom ≥ 2× the current step's traffic).
  • Owners: SRE Lead (model), Generation Platform Lead (GPU inputs), Realtime Lead (Psyche inputs), Finance Lead (unit economics).