# V6 Token-Cost Contingency Plan — Price Shocks, Outages, Kill Criteria

**Status:** Planning gap-fill per `V1_V7_PLAN_SET_AUDIT_2026-06-12.md` §6.2 ("no
model-price-shock or provider-outage contingency for a product whose stated
make-or-break bet is token cost"). **Date:** 2026-06-12. **Owners:** Moirai
kernel owner (tiering levers), model-ops lead (routing, caching, dashboard), V6
product lead (player-visible degradation decisions, kill-criteria escalation),
finance partner (weekly cost readings).

Canonical inputs: per-tier token budgets and concurrency caps in
`V6_DEPENDENCIES.md` §19 (:400–437) — Clotho ≤ 50k tokens/agent/active-minute
(typical 15–25k with caching) on Opus-class; Lachesis ≤ 10k/agent/reflection (≈
≤1,200/agent/game-minute amortised) on Sonnet-class; Atropos ≤ 15k/agent/
game-day and Vac ≤ 2k/parse on Haiku-class; Clotho scene cap design 16 / hard
24; Lachesis resident world 150–400; Solo-world hourly cognition cap with
"homestead rest". The dependency doc deliberately leaves the dollar figure
floating ("the dollar figure floats with provider pricing", dep:404–406); this
document is where the dollars get computed and stress-tested.

---

## 1. Reference prices and unit-cost model

Anthropic list prices (per MTok, as published in the Claude platform docs,
cached 2026-06-04; re-verify quarterly):

| Class (V6 routing, dep:426–428) | Model tier | Input | Output | Cache read (≈0.1× in) | Cache write 5-min (1.25× in) | Batch discount                          |
| ------------------------------- | ---------- | ----- | ------ | --------------------- | ---------------------------- | --------------------------------------- |
| Opus-class (Clotho)             | Opus       | $5.00 | $25.00 | $0.50                 | $6.25                        | n/a (interactive)                       |
| Sonnet-class (Lachesis)         | Sonnet     | $3.00 | $15.00 | $0.30                 | $3.75                        | −50% on everything                      |
| Haiku-class (Atropos, Vac)      | Haiku      | $1.00 | $5.00  | $0.10                 | $1.25                        | −50% (Atropos only; Vac is interactive) |

### 1.1 Reference scenario — cost per player-active-hour (all mix figures are planning assumptions adopted 2026-06-12, to be replaced by staging telemetry per the §19 cost gate)

**Clotho.** Full-cognition agent-minute at the _typical_ 20k tokens (mid of the
doc's 15–25k): split 90% input / 10% output ⇒ 18k in / 2k out; input mix 70%
cache-read / 10% cache-write / 20% fresh:

```
reads  12,600 × $0.50/M = $0.0063      writes 1,800 × $6.25/M = $0.0113
fresh   3,600 × $5.00/M = $0.0180      output 2,000 × $25.0/M = $0.0500
                              per full-cognition agent-minute ≈ $0.0856
```

Player attention is serial: assume dialogue/active-direction occupies 30% of
active minutes with on average 1.5 agents in full cognition ⇒ 27 full-cognition
agent-minutes per player-hour ⇒ $2.31. Ambient co-present Clotho agents (≤16
scene, plans cache-served, "cognition recomputed only on material change",
arch:1156–1158): 12 agents × 42 min × 0.8k tokens, ~90% cache-read input, sparse
output ⇒ ≈ $0.58. **Clotho ≈ $2.89/h** (input-side $1.34, output-side $1.55).

**Lachesis.** 250 resident agents (mid of 150–400); 40% have a due reflection
per cycle (the rest resolve from cached plans); 5 reflections/h (12-game-min
cadence, 1 game-min = 1 real-min online); typical reflection 4k tokens (vs the
10k cap), 85/15 in/out ⇒ 2.0M tokens/h. Batched Sonnet with 50% of input as
cache reads:

```
in  0.85M × $0.15 + 0.85M × $1.50 = $1.40      out 0.30M × $7.50 = $2.25
                                            Lachesis ≈ $3.65/h
```

**Atropos.** 500 distant/offline agents advancing per player at typical
6k/game-day (cap 15k), 1 game-day ≈ 2 real-h online ⇒ 1.5M tokens/h, batched
Haiku, 50% input cached ⇒ in $0.35 + out $0.56 ⇒ **≈ $0.91/h**.

**Vac.** 30 voice parses/h × 1.2k typical (cap 2k) ⇒ **≈ $0.05/h**. **Clio.**
Chronicle/arc-watch amortisation ⇒ **≈ $0.50/h** (0.30 in / 0.20 out).

| Tier                           | $/player-active-hour | Input-side | Output-side |
| ------------------------------ | -------------------- | ---------- | ----------- |
| Clotho (Opus-class)            | 2.89                 | 1.34       | 1.55        |
| Lachesis (Sonnet-class, batch) | 3.65                 | 1.40       | 2.25        |
| Atropos (Haiku-class, batch)   | 0.91                 | 0.35       | 0.56        |
| Vac (Haiku-class)              | 0.05                 | 0.03       | 0.02        |
| Clio (mixed, batch)            | 0.50                 | 0.30       | 0.20        |
| **Reference total**            | **$8.00**            | **$3.42**  | **$4.58**   |

Two readings of this table matter:

1. **The headline risk is not the Opus dialogue minutes — it is the 250 quietly
   reflecting residents.** Lachesis is the largest line despite the cheaper
   model, which is why the existing **[P2]** distillation note
   (V6_DEPENDENCIES.md:336) is the single biggest lever in §3.
2. **Budget caps ≠ expected spend.** If every tier ran at its cap (16 × 50k × 60
   Clotho; 400 × 1,200 × 60 Lachesis; Atropos at 15k), the ceiling is ≈ $205 +
   $53 + $2 ≈ **$260/player-active-hour** — two orders of magnitude above
   reference. The token budgets are correctly framed as the _contract_
   (dep:431–434); the dashboard target below is the _economic_ control,
   measured, not derived from caps.

**Dashboard target (planning assumption adopted 2026-06-12):** launch at the
$8.00 reference, trending to ≤ $5.00 by GA+2 quarters via §3 rungs 1–2. The
target is recorded on the model-ops cost dashboard the dependency doc already
designates as the live tracker.

---

## 2. Model-price-shock scenarios

The audit asks specifically about **input-price** moves; cache read/write prices
scale with input price (they are multipliers of it), so they shock together.
Output prices held constant in S1/S2.

| Scenario                            | Input-side          | Total $/player-active-hour | Delta  |
| ----------------------------------- | ------------------- | -------------------------- | ------ |
| S0 — today                          | $3.42               | **$8.00**                  | —      |
| S1 — input +50%                     | $3.42 × 1.5 = $5.13 | **$9.71**                  | +21.4% |
| S2 — input +100%                    | $3.42 × 2.0 = $6.84 | **$11.42**                 | +42.8% |
| S3 — full reprice +50% in _and_ out | —                   | **$12.00**                 | +50%   |
| S4 — full reprice +100% in and out  | —                   | **$16.00**                 | +100%  |

Observations: (a) V6 is _output-heavy_ ($4.58 of $8.00) because agents generate
dialogue, reflections, and beats — input-only shocks hurt less than intuition
suggests, and **output-token discipline (shorter reflections, schema-constrained
beats) is a real lever, not just caching**; (b) batch-tier work
(Lachesis+Atropos+Clio ≈ $5.06 of $8.00) reprices with the batch discount — if a
provider shock ever _removed_ the 50% batch discount instead, the hit would be
+$5.06/h (+63%), worse than S2; the contingency below treats "batch-discount
loss" as equivalent to S2-and-a-half and triggers the same ladder.

---

## 3. Mitigation ladder — in order, each rung quantified against S0

Rungs are ordered cheapest-first in player-experience cost; each is
independently deployable and gated by the existing eval posture (a routing or
cap change ships only with `V6/evals/*` suites green at their
`minimumPassRate`s, per arch:1307–1327).

| Rung  | Lever                                                                                            | Mechanics                                                                                                                                                                                                                                                                                          | Δ vs S0         | Running total |
| ----- | ------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------- | ------------- |
| **1** | **Cache hit-rate improvements** (current Clotho typical is already "15–25k _with caching_")      | Deterministic prompt assembly per shared-prefix rules; persona/policy/Ori-dossier prefixes on 1-h TTL caches keyed per agent; plan-cache reuse on the Ori (arch:1156–1158). Clotho input mix 70→85% reads (full-cog input $0.0356→$0.0223/agent-min); Lachesis batch reads 50→75% (in $1.40→$0.83) | **−$0.98**      | $7.02         |
| **2** | **Lachesis/Atropos distillation** — the existing P2 note (dep:336) promoted to P1 upon any shock | Distilled/fine-tuned Haiku-class for reflection and summary workloads; Lachesis post-rung-1 $3.08 → $1.03; Atropos −30%                                                                                                                                                                            | **−$2.32**      | $4.70         |
| **3** | **Scene-cap tightening 16 → 8** (design target; hard cap 24 → 12)                                | Halves ambient Clotho upkeep (−$0.27) and, more importantly, halves the _tail_: cap-spend ceiling $205 → $103/h; p95 scenes are where Clotho blowouts live                                                                                                                                         | **−$0.30 mean** | $4.40         |
| **4** | **Solo-world hourly-cap reduction + homestead rest default-on**                                  | Lachesis due-reflection fraction 40% → 27% (−$0.33 post-rung-2); Atropos batch −33% (−$0.21); the cap and rest multiplier already exist (dep:423–424, arch:1150–1155)                                                                                                                              | **−$0.54**      | **$3.86**     |

**Net:** the full ladder reaches ≈ **$3.86/h (−52%)**. Under S2 (+100% input)
applied to the post-ladder mix, cost ≈ $5.4/h — still **below today's
unmitigated $8.00**. The ladder absorbs a full +100% input-price shock; under S4
(full +100% reprice) post-ladder cost ≈ $7.7/h, i.e., S4 consumes the entire
ladder. Anything beyond S4 forces §5 structural moves.

Deliberately _not_ on the ladder: silently faking cognition (hardcoded
"reflections"), shrinking safety/eval coverage, or cutting the Isis gate —
forbidden under the quality standard; degradation must be honest (arch:
1167–1169: "the player is told honestly if their world is running degraded").

---

## 4. Provider-outage degradation — player-visible contract and recovery

The BT/HTN fallback is already specced: cheap execution is co-located with the
world server and survives a cognition outage (arch:657–660, 704, 728, 759–760);
"a cognition outage costs richness, never the world" (arch:422–424). This
section defines the _contract_ — what the player sees, tier by tier.

### Degradation tiers

| State                                            | Trigger                                  | Player-visible contract                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| ------------------------------------------------ | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **D0 normal**                                    | —                                        | Full fidelity.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| **D1 — Opus-class unavailable or 529-saturated** | Provider incident on the Clotho route    | Clotho reroutes to Sonnet-class. Honest banner: "Your companions are thinking a little more simply right now." Dialogue continues; behavior evals for the Sonnet route are pre-certified in CI (a standing `clotho-sonnet-fallback` eval profile kept green at all times, so the reroute is push-button, not a scramble).                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **D2 — all interactive cognition unavailable**   | Full provider outage / network partition | World never pauses: agents keep schedules, navigation, gestures, and cached barks via BT/HTN. Free conversation is declined honestly in-fiction _and_ in-UI ("Abeni can't find her words right now — she'll remember what you did"). Vac voice→intent parsing is down, but the structured intent builder still issues objectives (non-voice parity, arch:864–870) executed by BT/HTN. **No Ori writes are fabricated**: perception-derived events buffer durably (see `ori-dr-and-compaction.md` §3.5); reflections/summaries defer. Crossroads holds extend; **irreversible transitions freeze** — no `Departed`/`Transcended`/`Died` resolution while cognition is degraded (the Ereshkigal state machines simply do not advance, arch:983–1024). |
| **D3 — batch tier also unavailable**             | Extended outage                          | As D2, plus Chronicle/Book generation pauses; on the player's next return Clio covers the gap from the buffered events ("while the threads were tangled…") — which is exactly its narrative-reconciliation job (arch:894–897).                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |

### Recovery contract

- **Playability RTO: 0** (the world never stopped).
- **Full-fidelity restore: ≤ 15 min** after provider recovery (tier reassignment
  is recomputed every tick; rehydration is the normal escalation path,
  arch:1171–1173).
- **Backlog drain:** deferred Lachesis/Atropos work drains via batch at a
  catch-up budget ≤ 1.5× the normal hourly tier budget until clear; drain SLO 6
  h. Catch-up never exceeds per-tier token caps — an outage must not become a
  cost incident.
- **No retroactive fabrication:** drained reflections process buffered episodes
  with their original timestamps; if the player witnessed degraded behavior,
  Clio writes one reconciliation beat, logged as such.
- Multi-provider hedging is **out of scope at launch** (V6 consumes V1 model-ops
  routing; adding a second provider is a V1-platform decision). This plan's
  posture: degrade honestly within one provider, and keep the D1 Sonnet
  profile + distilled-model rung warm so the dependency is on _a_ capable model,
  not on one price point.

---

## 5. Kill criteria

Measured quantity: **blended cognition cost per player-active-hour**, weekly
reading from the model-ops dashboard (the metric dep:431–434 already tracks),
trailing-4-week average, evaluated every Monday by model-ops lead + finance
partner; product lead owns escalation.

**Trigger (X, Y):** trailing cost **> $12.00** (1.5× the $8.00 reference) for
**6 consecutive weekly readings** _with ladder rungs 1–2 already deployed_.

**Then (Z), staged:**

1. **Within 7 days:** deploy rungs 3–4 (scene cap 8, Solo hourly cap −33%,
   homestead rest default-on) with the §4 honesty banner and a player comms
   note. These are player-visible; shipping them is a product decision made
   _now_, in this document, so the on-call isn't negotiating it mid-incident.
2. **Within 30 days:** re-route Clotho default Opus-class → Sonnet-class behind
   the full agent-behavior + safety eval gates (all `minimumPassRate: 1` suites
   must stay at 1; `value-refusal`, `persona`, `crisis`, `minor-protection` are
   non-negotiable). If the Sonnet route cannot hold the eval bar, this step is
   **rejected** and we proceed to 3.
3. **If cost remains > $12.00 for 8 further weeks, or step 2 fails evals:**
   executive go/no-go on V6 live-service posture — halt regional expansion,
   freeze Commons cognition growth (interest-management radius down, wild- agent
   foundry seeding paused), and a formal decision memo on continuing, re-scoping
   (Solo-only product), or sunsetting, with the unit economics in this document
   attached.

**Emergency override:** trailing cost **> $20.00** (2.5×) for **2 consecutive
weeks** at any time ⇒ rungs 3–4 plus Clotho hard cap 8 immediately, without
waiting for the 6-week window.

**Standing review:** this plan's prices (§1) re-verified quarterly against the
provider price list; the reference model re-derived from staging telemetry at
every release via `verify:v6 moirai-cost-load-readiness`
(`V6/release/moirai-cost-load-readiness.v6release.json`), which remains the
binding release gate. All scenario arithmetic above is reproducible from the
tables in §1 — anyone disputing a number should be able to re-derive it on one
page, which is the point.
