Status: Planning gap-fill per V1_V7_PLAN_SET_AUDIT_2026-06-12.md §6.2 ("no
model-price-shock or provider-outage contingency for a product whose stated
make-or-break bet is token cost"). Date: 2026-06-12. Owners: Moirai
kernel owner (tiering levers), model-ops lead (routing, caching, dashboard), V6
product lead (player-visible degradation decisions, kill-criteria escalation),
finance partner (weekly cost readings).
Canonical inputs: per-tier token budgets and concurrency caps in
V6_DEPENDENCIES.md §19 (:400–437) — Clotho ≤ 50k tokens/agent/active-minute
(typical 15–25k with caching) on Opus-class; Lachesis ≤ 10k/agent/reflection (≈
≤1,200/agent/game-minute amortised) on Sonnet-class; Atropos ≤ 15k/agent/
game-day and Vac ≤ 2k/parse on Haiku-class; Clotho scene cap design 16 / hard
24; Lachesis resident world 150–400; Solo-world hourly cognition cap with
"homestead rest". The dependency doc deliberately leaves the dollar figure
floating ("the dollar figure floats with provider pricing", dep:404–406); this
document is where the dollars get computed and stress-tested.
1. Reference prices and unit-cost model#
Anthropic list prices (per MTok, as published in the Claude platform docs, cached 2026-06-04; re-verify quarterly):
| Class (V6 routing, dep:426–428) | Model tier | Input | Output | Cache read (≈0.1× in) | Cache write 5-min (1.25× in) | Batch discount |
|---|---|---|---|---|---|---|
| Opus-class (Clotho) | Opus | $5.00 | $25.00 | $0.50 | $6.25 | n/a (interactive) |
| Sonnet-class (Lachesis) | Sonnet | $3.00 | $15.00 | $0.30 | $3.75 | −50% on everything |
| Haiku-class (Atropos, Vac) | Haiku | $1.00 | $5.00 | $0.10 | $1.25 | −50% (Atropos only; Vac is interactive) |
1.1 Reference scenario — cost per player-active-hour (all mix figures are planning assumptions adopted 2026-06-12, to be replaced by staging telemetry per the §19 cost gate)#
Clotho. Full-cognition agent-minute at the typical 20k tokens (mid of the doc's 15–25k): split 90% input / 10% output ⇒ 18k in / 2k out; input mix 70% cache-read / 10% cache-write / 20% fresh:
reads 12,600 × $0.50/M = $0.0063 writes 1,800 × $6.25/M = $0.0113
fresh 3,600 × $5.00/M = $0.0180 output 2,000 × $25.0/M = $0.0500
per full-cognition agent-minute ≈ $0.0856
Player attention is serial: assume dialogue/active-direction occupies 30% of active minutes with on average 1.5 agents in full cognition ⇒ 27 full-cognition agent-minutes per player-hour ⇒ $2.31. Ambient co-present Clotho agents (≤16 scene, plans cache-served, "cognition recomputed only on material change", arch:1156–1158): 12 agents × 42 min × 0.8k tokens, ~90% cache-read input, sparse output ⇒ ≈ $0.58. Clotho ≈ $2.89/h (input-side $1.34, output-side $1.55).
Lachesis. 250 resident agents (mid of 150–400); 40% have a due reflection per cycle (the rest resolve from cached plans); 5 reflections/h (12-game-min cadence, 1 game-min = 1 real-min online); typical reflection 4k tokens (vs the 10k cap), 85/15 in/out ⇒ 2.0M tokens/h. Batched Sonnet with 50% of input as cache reads:
in 0.85M × $0.15 + 0.85M × $1.50 = $1.40 out 0.30M × $7.50 = $2.25
Lachesis ≈ $3.65/h
Atropos. 500 distant/offline agents advancing per player at typical 6k/game-day (cap 15k), 1 game-day ≈ 2 real-h online ⇒ 1.5M tokens/h, batched Haiku, 50% input cached ⇒ in $0.35 + out $0.56 ⇒ ≈ $0.91/h.
Vac. 30 voice parses/h × 1.2k typical (cap 2k) ⇒ ≈ $0.05/h. Clio. Chronicle/arc-watch amortisation ⇒ ≈ $0.50/h (0.30 in / 0.20 out).
| Tier | $/player-active-hour | Input-side | Output-side |
|---|---|---|---|
| Clotho (Opus-class) | 2.89 | 1.34 | 1.55 |
| Lachesis (Sonnet-class, batch) | 3.65 | 1.40 | 2.25 |
| Atropos (Haiku-class, batch) | 0.91 | 0.35 | 0.56 |
| Vac (Haiku-class) | 0.05 | 0.03 | 0.02 |
| Clio (mixed, batch) | 0.50 | 0.30 | 0.20 |
| Reference total | $8.00 | $3.42 | $4.58 |
Two readings of this table matter:
- The headline risk is not the Opus dialogue minutes — it is the 250 quietly reflecting residents. Lachesis is the largest line despite the cheaper model, which is why the existing [P2] distillation note (V6_DEPENDENCIES.md:336) is the single biggest lever in §3.
- Budget caps ≠ expected spend. If every tier ran at its cap (16 × 50k × 60 Clotho; 400 × 1,200 × 60 Lachesis; Atropos at 15k), the ceiling is ≈ $205 + $53 + $2 ≈ $260/player-active-hour — two orders of magnitude above reference. The token budgets are correctly framed as the contract (dep:431–434); the dashboard target below is the economic control, measured, not derived from caps.
Dashboard target (planning assumption adopted 2026-06-12): launch at the $8.00 reference, trending to ≤ $5.00 by GA+2 quarters via §3 rungs 1–2. The target is recorded on the model-ops cost dashboard the dependency doc already designates as the live tracker.
2. Model-price-shock scenarios#
The audit asks specifically about input-price moves; cache read/write prices scale with input price (they are multipliers of it), so they shock together. Output prices held constant in S1/S2.
| Scenario | Input-side | Total $/player-active-hour | Delta |
|---|---|---|---|
| S0 — today | $3.42 | $8.00 | — |
| S1 — input +50% | $3.42 × 1.5 = $5.13 | $9.71 | +21.4% |
| S2 — input +100% | $3.42 × 2.0 = $6.84 | $11.42 | +42.8% |
| S3 — full reprice +50% in and out | — | $12.00 | +50% |
| S4 — full reprice +100% in and out | — | $16.00 | +100% |
Observations: (a) V6 is output-heavy ($4.58 of $8.00) because agents generate dialogue, reflections, and beats — input-only shocks hurt less than intuition suggests, and output-token discipline (shorter reflections, schema-constrained beats) is a real lever, not just caching; (b) batch-tier work (Lachesis+Atropos+Clio ≈ $5.06 of $8.00) reprices with the batch discount — if a provider shock ever removed the 50% batch discount instead, the hit would be +$5.06/h (+63%), worse than S2; the contingency below treats "batch-discount loss" as equivalent to S2-and-a-half and triggers the same ladder.
3. Mitigation ladder — in order, each rung quantified against S0#
Rungs are ordered cheapest-first in player-experience cost; each is
independently deployable and gated by the existing eval posture (a routing or
cap change ships only with V6/evals/* suites green at their
minimumPassRates, per arch:1307–1327).
| Rung | Lever | Mechanics | Δ vs S0 | Running total |
|---|---|---|---|---|
| 1 | Cache hit-rate improvements (current Clotho typical is already "15–25k with caching") | Deterministic prompt assembly per shared-prefix rules; persona/policy/Ori-dossier prefixes on 1-h TTL caches keyed per agent; plan-cache reuse on the Ori (arch:1156–1158). Clotho input mix 70→85% reads (full-cog input $0.0356→$0.0223/agent-min); Lachesis batch reads 50→75% (in $1.40→$0.83) | −$0.98 | $7.02 |
| 2 | Lachesis/Atropos distillation — the existing P2 note (dep:336) promoted to P1 upon any shock | Distilled/fine-tuned Haiku-class for reflection and summary workloads; Lachesis post-rung-1 $3.08 → $1.03; Atropos −30% | −$2.32 | $4.70 |
| 3 | Scene-cap tightening 16 → 8 (design target; hard cap 24 → 12) | Halves ambient Clotho upkeep (−$0.27) and, more importantly, halves the tail: cap-spend ceiling $205 → $103/h; p95 scenes are where Clotho blowouts live | −$0.30 mean | $4.40 |
| 4 | Solo-world hourly-cap reduction + homestead rest default-on | Lachesis due-reflection fraction 40% → 27% (−$0.33 post-rung-2); Atropos batch −33% (−$0.21); the cap and rest multiplier already exist (dep:423–424, arch:1150–1155) | −$0.54 | $3.86 |
Net: the full ladder reaches ≈ $3.86/h (−52%). Under S2 (+100% input) applied to the post-ladder mix, cost ≈ $5.4/h — still below today's unmitigated $8.00. The ladder absorbs a full +100% input-price shock; under S4 (full +100% reprice) post-ladder cost ≈ $7.7/h, i.e., S4 consumes the entire ladder. Anything beyond S4 forces §5 structural moves.
Deliberately not on the ladder: silently faking cognition (hardcoded "reflections"), shrinking safety/eval coverage, or cutting the Isis gate — forbidden under the quality standard; degradation must be honest (arch: 1167–1169: "the player is told honestly if their world is running degraded").
4. Provider-outage degradation — player-visible contract and recovery#
The BT/HTN fallback is already specced: cheap execution is co-located with the world server and survives a cognition outage (arch:657–660, 704, 728, 759–760); "a cognition outage costs richness, never the world" (arch:422–424). This section defines the contract — what the player sees, tier by tier.
Degradation tiers#
| State | Trigger | Player-visible contract |
|---|---|---|
| D0 normal | — | Full fidelity. |
| D1 — Opus-class unavailable or 529-saturated | Provider incident on the Clotho route | Clotho reroutes to Sonnet-class. Honest banner: "Your companions are thinking a little more simply right now." Dialogue continues; behavior evals for the Sonnet route are pre-certified in CI (a standing clotho-sonnet-fallback eval profile kept green at all times, so the reroute is push-button, not a scramble). |
| D2 — all interactive cognition unavailable | Full provider outage / network partition | World never pauses: agents keep schedules, navigation, gestures, and cached barks via BT/HTN. Free conversation is declined honestly in-fiction and in-UI ("Abeni can't find her words right now — she'll remember what you did"). Vac voice→intent parsing is down, but the structured intent builder still issues objectives (non-voice parity, arch:864–870) executed by BT/HTN. No Ori writes are fabricated: perception-derived events buffer durably (see ori-dr-and-compaction.md §3.5); reflections/summaries defer. Crossroads holds extend; irreversible transitions freeze — no Departed/Transcended/Died resolution while cognition is degraded (the Ereshkigal state machines simply do not advance, arch:983–1024). |
| D3 — batch tier also unavailable | Extended outage | As D2, plus Chronicle/Book generation pauses; on the player's next return Clio covers the gap from the buffered events ("while the threads were tangled…") — which is exactly its narrative-reconciliation job (arch:894–897). |
Recovery contract#
- Playability RTO: 0 (the world never stopped).
- Full-fidelity restore: ≤ 15 min after provider recovery (tier reassignment is recomputed every tick; rehydration is the normal escalation path, arch:1171–1173).
- Backlog drain: deferred Lachesis/Atropos work drains via batch at a catch-up budget ≤ 1.5× the normal hourly tier budget until clear; drain SLO 6 h. Catch-up never exceeds per-tier token caps — an outage must not become a cost incident.
- No retroactive fabrication: drained reflections process buffered episodes with their original timestamps; if the player witnessed degraded behavior, Clio writes one reconciliation beat, logged as such.
- Multi-provider hedging is out of scope at launch (V6 consumes V1 model-ops routing; adding a second provider is a V1-platform decision). This plan's posture: degrade honestly within one provider, and keep the D1 Sonnet profile + distilled-model rung warm so the dependency is on a capable model, not on one price point.
5. Kill criteria#
Measured quantity: blended cognition cost per player-active-hour, weekly reading from the model-ops dashboard (the metric dep:431–434 already tracks), trailing-4-week average, evaluated every Monday by model-ops lead + finance partner; product lead owns escalation.
Trigger (X, Y): trailing cost > $12.00 (1.5× the $8.00 reference) for 6 consecutive weekly readings with ladder rungs 1–2 already deployed.
Then (Z), staged:
- Within 7 days: deploy rungs 3–4 (scene cap 8, Solo hourly cap −33%, homestead rest default-on) with the §4 honesty banner and a player comms note. These are player-visible; shipping them is a product decision made now, in this document, so the on-call isn't negotiating it mid-incident.
- Within 30 days: re-route Clotho default Opus-class → Sonnet-class behind
the full agent-behavior + safety eval gates (all
minimumPassRate: 1suites must stay at 1;value-refusal,persona,crisis,minor-protectionare non-negotiable). If the Sonnet route cannot hold the eval bar, this step is rejected and we proceed to 3. - If cost remains > $12.00 for 8 further weeks, or step 2 fails evals: executive go/no-go on V6 live-service posture — halt regional expansion, freeze Commons cognition growth (interest-management radius down, wild- agent foundry seeding paused), and a formal decision memo on continuing, re-scoping (Solo-only product), or sunsetting, with the unit economics in this document attached.
Emergency override: trailing cost > $20.00 (2.5×) for 2 consecutive weeks at any time ⇒ rungs 3–4 plus Clotho hard cap 8 immediately, without waiting for the 6-week window.
Standing review: this plan's prices (§1) re-verified quarterly against the
provider price list; the reference model re-derived from staging telemetry at
every release via verify:v6 moirai-cost-load-readiness
(V6/release/moirai-cost-load-readiness.v6release.json), which remains the
binding release gate. All scenario arithmetic above is reproducible from the
tables in §1 — anyone disputing a number should be able to re-derive it on one
page, which is the point.