# Eve SOTA embedding and reranker pricing — `eve.embedding-pricing.v1`

- **Task:** 3.2
- **Evaluated:** 2026-09-05
- **Decision basis:** pricing-and-posture
- **Chosen:** `perplexity/pplx-embed-v1-0.6b`
- **Registry leg:** `embedding` — bound
- **Probe spend:** $0.28906573 over 903 cost-reported requests
- **Record digest:**
  `86122d349529f9682e6f0f8b38c1939489c29d1e4203b24a97d7347e8fcdd219`

## What was measured, and why a catalogue was not enough

OpenRouter's default model listing omits embedding and rerank models entirely,
so the candidate set was ASKED for (`/models?output_modalities=embeddings`,
`/models?output_modalities=rerank`) rather than guessed: 37 embedding candidates
and 7 rerankers. Every price below is `usage.cost` returned by a completed call,
never `pricing.prompt` from the catalogue — one reranker in this table lists its
prompt price as `0` and bills a real amount.

The corpus is exact: 56,970 chunks, 36,147,926 characters, longest chunk 1205
characters (319 tokens at the widest tokenizer measured). Token volume is
projected by a ratio estimator over disjoint calibration batches, with its
interval.

## Admitted candidates, cheapest first

| Slug                                | Dims | $/1M    | Index build               | Per query | Query p50 | Endpoints | Caveats                                                                                                            |
| ----------------------------------- | ---- | ------- | ------------------------- | --------- | --------- | --------- | ------------------------------------------------------------------------------------------------------------------ |
| `perplexity/pplx-embed-v1-0.6b`     | 1024 | $0.0040 | $0.0359 ($0.0354–$0.0364) | $3.07e-8  | 203.5 ms  | 1         | single-endpoint-no-failover                                                                                        |
| `baai/bge-base-en-v1.5`             | 768  | $0.0050 | $0.0471 ($0.0462–$0.0479) | $5.17e-8  | 309.4 ms  | 1         | single-endpoint-no-failover, quantization-undeclared                                                               |
| `thenlper/gte-large`                | 1024 | $0.0100 | $0.0941 ($0.0924–$0.0959) | $1.03e-7  | 320.6 ms  | 1         | single-endpoint-no-failover, quantization-undeclared                                                               |
| `baai/bge-large-en-v1.5`            | 1024 | $0.0100 | $0.0941 ($0.0924–$0.0959) | $1.03e-7  | 325.3 ms  | 1         | single-endpoint-no-failover, quantization-undeclared                                                               |
| `baai/bge-m3`                       | 1024 | $0.0100 | $0.1066 ($0.1049–$0.1083) | $9.33e-8  | 338.9 ms  | 2         | —                                                                                                                  |
| `qwen/qwen3-embedding-8b`           | 4096 | $0.0100 | $0.0903 ($0.0890–$0.0916) | $8.67e-8  | 608.2 ms  | 3         | cross-endpoint-reordering, requires-documented-query-protocol                                                      |
| `voyageai/voyage-4-lite`            | 1024 | $0.0200 | $0.1795 ($0.1769–$0.1821) | $1.53e-7  | 244.4 ms  | 1         | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared                              |
| `openai/text-embedding-3-small`     | 1536 | $0.0200 | $0.1749 ($0.1722–$0.1776) | $1.40e-7  | 449.1 ms  | 2         | quantization-undeclared                                                                                            |
| `qwen/qwen3-embedding-4b`           | 2560 | $0.0200 | $0.1807 ($0.1780–$0.1833) | $1.73e-7  | 504.6 ms  | 1         | single-endpoint-no-failover, quantization-undeclared, requires-documented-query-protocol                           |
| `perplexity/pplx-embed-v1-4b`       | 2560 | $0.0300 | $0.2693 ($0.2654–$0.2731) | $2.30e-7  | 209.1 ms  | 1         | single-endpoint-no-failover                                                                                        |
| `voyageai/voyage-4`                 | 1024 | $0.0600 | $0.5385 ($0.5308–$0.5463) | $4.60e-7  | 212.3 ms  | 1         | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared                              |
| `mistralai/mistral-embed-2312`      | 1024 | $0.1000 | $1.0910 ($1.0769–$1.1050) | $1.07e-6  | 160.3 ms  | 3         | quantization-undeclared                                                                                            |
| `openai/text-embedding-ada-002`     | 1536 | $0.1000 | $0.8746 ($0.8612–$0.8880) | $7.00e-7  | 471.5 ms  | 1         | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared                              |
| `voyageai/voyage-multimodal-3.5`    | 1024 | $0.1200 | $1.0770 ($1.0616–$1.0925) | $9.20e-7  | 269.7 ms  | 1         | dimensions-silently-ignored, single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
| `voyageai/voyage-4-large`           | 1024 | $0.1200 | $1.0770 ($1.0616–$1.0925) | $9.20e-7  | 240.8 ms  | 1         | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared                              |
| `openai/text-embedding-3-large`     | 3072 | $0.1300 | $1.1370 ($1.1195–$1.1544) | $9.10e-7  | 513.8 ms  | 2         | cross-endpoint-reordering, quantization-undeclared                                                                 |
| `google/gemini-embedding-001`       | 3072 | $0.1500 | $1.4047 ($1.3785–$1.4309) | $1.15e-6  | 361.6 ms  | 2         | quantization-undeclared                                                                                            |
| `mistralai/codestral-embed-2505`    | 1536 | $0.1500 | $1.4203 ($1.4006–$1.4400) | $1.50e-6  | 201.4 ms  | 3         | quantization-undeclared                                                                                            |
| `google/gemini-embedding-2`         | 3072 | $0.2000 | $1.8729 ($1.8380–$1.9079) | $1.53e-6  | 413.3 ms  | 4         | quantization-undeclared                                                                                            |
| `google/gemini-embedding-2-preview` | 3072 | $0.2000 | $1.8729 ($1.8380–$1.9079) | $1.53e-6  | 335.3 ms  | 1         | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared                              |

20 of 37 candidates were admitted; 17 were refused.

## Refused, with the measurement that refused them

| Slug                                               | Refusals                                                                                                 | The measurement                                                                                  |
| -------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
| `liquid/lfm-2.5-embedding-350m:free`               | trains-on-input                                                                                          | truncation: refused-oversized-input; no endpoint accepts `data_collection: deny`                 |
| `voyageai/voyage-code-4`                           | identity-floor-failed, uncalibrated                                                                      | plain/nav-admin-repo-map first relevant at rank 6                                                |
| `nvidia/nemotron-3-embed-1b:free`                  | trains-on-input                                                                                          | no endpoint accepts `data_collection: deny`                                                      |
| `google/gemini-embedding-2:batch`                  | batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call                                         |
| `nvidia/llama-nemotron-embed-vl-1b-v2:free`        | trains-on-input                                                                                          | no endpoint accepts `data_collection: deny`                                                      |
| `thenlper/gte-base`                                | silently-truncates, uncalibrated                                                                         | truncation: token-count-plateaus at 512 tokens                                                   |
| `intfloat/e5-large-v2`                             | silently-truncates, uncalibrated                                                                         | truncation: token-count-plateaus at 512 tokens                                                   |
| `intfloat/e5-base-v2`                              | silently-truncates, uncalibrated                                                                         | truncation: token-count-plateaus at 512 tokens                                                   |
| `intfloat/multilingual-e5-large`                   | silently-truncates, uncalibrated                                                                         | truncation: token-count-plateaus at 512 tokens                                                   |
| `sentence-transformers/paraphrase-minilm-l6-v2`    | silently-truncates, context-below-corpus-maximum, uncalibrated                                           | truncation: token-count-plateaus at 128 tokens                                                   |
| `sentence-transformers/all-minilm-l12-v2`          | silently-truncates, identity-floor-failed, context-below-corpus-maximum, uncalibrated                    | truncation: token-count-plateaus at 128 tokens; plain/id-admin-adr-0076 first relevant at rank 8 |
| `sentence-transformers/multi-qa-mpnet-base-dot-v1` | silently-truncates, uncalibrated                                                                         | truncation: token-count-plateaus at 512 tokens                                                   |
| `sentence-transformers/all-mpnet-base-v2`          | silently-truncates, uncalibrated                                                                         | truncation: token-count-plateaus at 384 tokens                                                   |
| `sentence-transformers/all-minilm-l6-v2`           | silently-truncates, context-below-corpus-maximum, uncalibrated                                           | truncation: token-count-plateaus at 256 tokens                                                   |
| `openai/text-embedding-ada-002:batch`              | batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call                                         |
| `openai/text-embedding-3-large:batch`              | batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call                                         |
| `openai/text-embedding-3-small:batch`              | batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call                                         |

## The identity floor

Not a relevance measurement — task 3.5 owns those. This is the bar the cost rule
implies: the rule binds "the cheapest model that CAN DO THE JOB", so a price
ranking with no capability bar underneath it would rank models that cannot
retrieve at all. Each candidate must put a page named by its own identifier
above 200 chunks drawn from pages the label does not point at, using only the
corpus-identifier labels task 3.1 settled without any ranker.

Several of these models are asymmetric by design and publish a query-side
prefix. A model called without the one it documents is a model called wrong, so
each is measured under every protocol its family documents AND under plain text
as the control.

| Slug                                               | Protocol       | Result                                                                                                                                                             |
| -------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `liquid/lfm-2.5-embedding-350m:free`               | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2                                                                                                |
| `voyageai/voyage-code-4`                           | none cleared   | plain/nav-admin-repo-map: FAIL @6; plain/id-admin-adr-0076: pass @1                                                                                                |
| `voyageai/voyage-multimodal-3.5`                   | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `voyageai/voyage-4-lite`                           | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `voyageai/voyage-4`                                | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `voyageai/voyage-4-large`                          | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `nvidia/nemotron-3-embed-1b:free`                  | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `google/gemini-embedding-2`                        | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `google/gemini-embedding-2-preview`                | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `perplexity/pplx-embed-v1-4b`                      | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `perplexity/pplx-embed-v1-0.6b`                    | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3                                                                                                |
| `nvidia/llama-nemotron-embed-vl-1b-v2:free`        | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `thenlper/gte-base`                                | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `thenlper/gte-large`                               | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `intfloat/e5-large-v2`                             | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1     |
| `intfloat/e5-base-v2`                              | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1     |
| `intfloat/multilingual-e5-large`                   | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1     |
| `sentence-transformers/paraphrase-minilm-l6-v2`    | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3                                                                                                |
| `sentence-transformers/all-minilm-l12-v2`          | none cleared   | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @8                                                                                                |
| `baai/bge-base-en-v1.5`                            | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; bge-en-instruction/nav-admin-repo-map: pass @1; bge-en-instruction/id-admin-adr-0076: pass @1 |
| `sentence-transformers/multi-qa-mpnet-base-dot-v1` | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @4                                                                                                |
| `baai/bge-large-en-v1.5`                           | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; bge-en-instruction/nav-admin-repo-map: pass @1; bge-en-instruction/id-admin-adr-0076: pass @1 |
| `baai/bge-m3`                                      | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `sentence-transformers/all-mpnet-base-v2`          | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2                                                                                                |
| `sentence-transformers/all-minilm-l6-v2`           | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3                                                                                                |
| `mistralai/mistral-embed-2312`                     | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `google/gemini-embedding-001`                      | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `openai/text-embedding-ada-002`                    | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3                                                                                                |
| `mistralai/codestral-embed-2505`                   | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1                                                                                                |
| `openai/text-embedding-3-large`                    | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2                                                                                                |
| `openai/text-embedding-3-small`                    | plain          | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2                                                                                                |
| `qwen/qwen3-embedding-8b`                          | qwen3-instruct | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @7; qwen3-instruct/nav-admin-repo-map: pass @1; qwen3-instruct/id-admin-adr-0076: pass @2         |
| `qwen/qwen3-embedding-4b`                          | qwen3-instruct | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @16; qwen3-instruct/nav-admin-repo-map: pass @1; qwen3-instruct/id-admin-adr-0076: pass @1        |

The floor uses 2 corpus-identifier queries against 200 distractor chunks. Its
protocols: `plain` (no prefix; the control every candidate is measured under);
`e5-query-passage` (intfloat E5 family model cards: "query: " and "passage: "
prefixes); `bge-en-instruction` (BAAI bge-\*-en-v1.5 model cards: a query-side
retrieval instruction); `qwen3-instruct` (Qwen3-Embedding model cards: an
instruction-templated query side).

## Provider data posture

Posture is measured by asking the router for it. A request that routes had an
endpoint meeting the filter; one that 404s had none, which the router could only
know by checking every endpoint.

| Slug                                               | Baseline | Zero data retention | Training denied |
| -------------------------------------------------- | -------- | ------------------- | --------------- |
| `liquid/lfm-2.5-embedding-350m:free`               | routed   | refused             | refused         |
| `voyageai/voyage-code-4`                           | routed   | refused             | routed          |
| `voyageai/voyage-multimodal-3.5`                   | routed   | refused             | routed          |
| `voyageai/voyage-4-lite`                           | routed   | refused             | routed          |
| `voyageai/voyage-4`                                | routed   | refused             | routed          |
| `voyageai/voyage-4-large`                          | routed   | refused             | routed          |
| `nvidia/nemotron-3-embed-1b:free`                  | routed   | refused             | refused         |
| `google/gemini-embedding-2`                        | routed   | routed              | routed          |
| `google/gemini-embedding-2-preview`                | routed   | refused             | routed          |
| `perplexity/pplx-embed-v1-4b`                      | routed   | routed              | routed          |
| `perplexity/pplx-embed-v1-0.6b`                    | routed   | routed              | routed          |
| `nvidia/llama-nemotron-embed-vl-1b-v2:free`        | routed   | refused             | refused         |
| `thenlper/gte-base`                                | routed   | routed              | routed          |
| `thenlper/gte-large`                               | routed   | routed              | routed          |
| `intfloat/e5-large-v2`                             | routed   | routed              | routed          |
| `intfloat/e5-base-v2`                              | routed   | routed              | routed          |
| `intfloat/multilingual-e5-large`                   | routed   | routed              | routed          |
| `sentence-transformers/paraphrase-minilm-l6-v2`    | routed   | routed              | routed          |
| `sentence-transformers/all-minilm-l12-v2`          | routed   | routed              | routed          |
| `baai/bge-base-en-v1.5`                            | routed   | routed              | routed          |
| `sentence-transformers/multi-qa-mpnet-base-dot-v1` | routed   | routed              | routed          |
| `baai/bge-large-en-v1.5`                           | routed   | routed              | routed          |
| `baai/bge-m3`                                      | routed   | routed              | routed          |
| `sentence-transformers/all-mpnet-base-v2`          | routed   | routed              | routed          |
| `sentence-transformers/all-minilm-l6-v2`           | routed   | routed              | routed          |
| `mistralai/mistral-embed-2312`                     | routed   | routed              | routed          |
| `google/gemini-embedding-001`                      | routed   | routed              | routed          |
| `openai/text-embedding-ada-002`                    | routed   | refused             | routed          |
| `mistralai/codestral-embed-2505`                   | routed   | routed              | routed          |
| `openai/text-embedding-3-large`                    | routed   | routed              | routed          |
| `openai/text-embedding-3-small`                    | routed   | routed              | routed          |
| `qwen/qwen3-embedding-8b`                          | routed   | routed              | routed          |
| `qwen/qwen3-embedding-4b`                          | routed   | routed              | routed          |

## Reranking, priced against what it would be added to

| Slug                                         | Served by           | Billing unit | $/query     | Latency  | × the embedding query cost |
| -------------------------------------------- | ------------------- | ------------ | ----------- | -------- | -------------------------- |
| `qwen/qwen3-reranker-8b`                     | Fireworks           | unknown      | $0.0022886  | 567.2 ms | 74628.26×                  |
| `voyageai/rerank-2.5-lite`                   | VoyageAI by MongoDB | unknown      | $0.00015334 | 389.8 ms | 5000.22×                   |
| `voyageai/rerank-2.5`                        | VoyageAI by MongoDB | unknown      | $0.00038335 | 358.5 ms | 12500.54×                  |
| `nvidia/llama-nemotron-rerank-vl-1b-v2:free` | Nvidia              | unknown      | $0          | 729.3 ms | 0×                         |
| `cohere/rerank-4-pro`                        | Cohere              | search-unit  | $0.0025     | 521 ms   | 81521.74×                  |
| `cohere/rerank-4-fast`                       | Cohere              | search-unit  | $0.002      | 366.5 ms | 65217.39×                  |
| `cohere/rerank-v3.5`                         | Cohere              | search-unit  | $0.001      | 279.6 ms | 32608.7×                   |

deferred to task 3.4 — priced here and not adopted; the measured per-query cost
is the reason it needs a quality argument before it is added, not a footnote.

One row bills nothing, and a zero there is a price, not a posture: every `:free`
route in the embedding field was refused because no endpoint of it accepts
`data_collection: deny`, and the refusal named "Free model training". A free
reranker earns the same check before it is adopted.

## The binding

`perplexity/pplx-embed-v1-0.6b` is bound because it is the cheapest candidate
that cleared every measured gate at 10,000 queries/month — $0.0040 per million
tokens, 1024 dimensions, served by Perplexity. The runner-up is
`baai/bge-base-en-v1.5`, $0.3352 per month dearer at that volume.

The binding is the slug, the endpoint set AND the query protocol `plain` — an
index built under one convention and queried under another is two systems, not
one.

An endpoint pin is still applied: this slug has a single endpoint, so no
agreement between endpoints was measured and none is claimed; the pin names the
endpoint that was.

### Rollback

- Slug override: `OSHUN_ASSISTANT_EMBEDDING_MODEL` (registry supplies the
  default underneath it)
- Endpoint override: `OPENROUTER_PROVIDER_ONLY`
- Degraded mode: `lexical-only` — apps/oshun/bff/src/assistant/docs-search.ts
- Rolling back invalidates stored vectors: true

### What task 3.5 must measure

Task 3.2 binds on price alone, and a price-only pin handed to 3.5 would leave it
measuring one arm chosen by cost. The shortlist below is the admitted
price/width frontier — for every admitted candidate NOT on it, some admitted
candidate is at least as cheap and at least as wide. Dimensions are a capacity
fact here, not a claim about retrieval.

- `perplexity/pplx-embed-v1-0.6b`
- `qwen/qwen3-embedding-8b`

### What re-opens this binding

- task 3.5 measures relevance and a different candidate wins on a paired test
- task 3.7 declines to promote dense retrieval at all, which retires this leg
- the bound endpoint set changes, because vectors are only comparable within one
  endpoint
- a re-priced route moves the chosen candidate out of first place at the
  recorded volume
- the corpus chunker changes, because the calibrated tokens-per-character no
  longer applies
- the query-side protocol changes, because an index built under one convention
  cannot be queried under another
- the bound slug gains or loses an endpoint whose quantization differs from the
  measured one
- the vendor re-points a versioned slug, because a vendor-versioned name is a
  pin only while the vendor keeps it one

## Honest limits

- This binds on PRICE and POSTURE. No relevance measurement exists yet: task 3.5
  owns Recall@k and nDCG over the full set with a pre-registered paired test,
  and task 3.7 owns promotion. A candidate that clears the identity floor has
  been shown able to retrieve, not shown to retrieve well.
- The identity floor uses only the relevance set's corpus-identifier labels,
  because those are the ones that settle without any ranker. Four of the nine
  query classes still carry no labels at all (task 3.1), so no floor exists for
  conceptual, multi-hop, recency or contradictory queries.
- Corpus token counts are a ratio estimator over disjoint sampled batches, not
  an exact count. The corpus CHARACTER count is exact; buying an exact token
  count would mean embedding all 36 million characters for every candidate.
- Monthly figures are scenarios. This repository records no `search_docs` call
  rate, so the per-query and per-index prices are the measured facts and the
  monthly totals are arithmetic over a volume a reader chooses.
- Latency is wall-clock from this host on this date, over a residential-grade
  path to one region. It is a floor for planning, not a service-level
  measurement.
- Posture is measured as ROUTING behaviour — whether an endpoint meeting a
  filter exists — not as an audit of any provider's actual practice.
- Nothing is indexed. Task 3.3 builds the index, and until it does the bound
  slug has no production consumer; `search_docs` remains lexical BM25F.
- Prices are quoted for text as the corpus stores it. A binding whose query
  protocol adds a prefix costs a few tokens more per query and, where the
  protocol also prefixes documents, roughly two per cent more per index build;
  neither shifts the ranking between candidates at these magnitudes, and the
  protocol is recorded so the difference is recomputable.
- Each price is the single `usage.cost` figure the route returned for one call,
  retained verbatim in the probe record beside its token count. This record
  therefore distinguishes no cached from uncached rate, and nothing here should
  be read as a claim that one exists or does not.
- The identity floor is measured under the protocols this tool knows about. A
  model documenting a calling convention nobody here recognised is measured
  plain and would be under-measured exactly as the instruction-tuned models were
  before their protocols were added.
