Disciplines · Audits

Eve SOTA embedding and reranker pricing — eve.embedding-pricing.v1

OpenRouter's default model listing omits embedding and rerank models entirely, so the candidate set was ASKED for (/models?output_modalities=embeddings,

8sections15 minread

On this page
  • Task: 3.2
  • Evaluated: 2026-09-05
  • Decision basis: pricing-and-posture
  • Chosen: perplexity/pplx-embed-v1-0.6b
  • Registry leg: embedding — bound
  • Probe spend: $0.28906573 over 903 cost-reported requests
  • Record digest: 86122d349529f9682e6f0f8b38c1939489c29d1e4203b24a97d7347e8fcdd219

What was measured, and why a catalogue was not enough#

OpenRouter's default model listing omits embedding and rerank models entirely, so the candidate set was ASKED for (/models?output_modalities=embeddings, /models?output_modalities=rerank) rather than guessed: 37 embedding candidates and 7 rerankers. Every price below is usage.cost returned by a completed call, never pricing.prompt from the catalogue — one reranker in this table lists its prompt price as 0 and bills a real amount.

The corpus is exact: 56,970 chunks, 36,147,926 characters, longest chunk 1205 characters (319 tokens at the widest tokenizer measured). Token volume is projected by a ratio estimator over disjoint calibration batches, with its interval.

Admitted candidates, cheapest first#

Slug Dims $/1M Index build Per query Query p50 Endpoints Caveats
perplexity/pplx-embed-v1-0.6b 1024 $0.0040 $0.0359 ($0.0354–$0.0364) $3.07e-8 203.5 ms 1 single-endpoint-no-failover
baai/bge-base-en-v1.5 768 $0.0050 $0.0471 ($0.0462–$0.0479) $5.17e-8 309.4 ms 1 single-endpoint-no-failover, quantization-undeclared
thenlper/gte-large 1024 $0.0100 $0.0941 ($0.0924–$0.0959) $1.03e-7 320.6 ms 1 single-endpoint-no-failover, quantization-undeclared
baai/bge-large-en-v1.5 1024 $0.0100 $0.0941 ($0.0924–$0.0959) $1.03e-7 325.3 ms 1 single-endpoint-no-failover, quantization-undeclared
baai/bge-m3 1024 $0.0100 $0.1066 ($0.1049–$0.1083) $9.33e-8 338.9 ms 2
qwen/qwen3-embedding-8b 4096 $0.0100 $0.0903 ($0.0890–$0.0916) $8.67e-8 608.2 ms 3 cross-endpoint-reordering, requires-documented-query-protocol
voyageai/voyage-4-lite 1024 $0.0200 $0.1795 ($0.1769–$0.1821) $1.53e-7 244.4 ms 1 single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared
openai/text-embedding-3-small 1536 $0.0200 $0.1749 ($0.1722–$0.1776) $1.40e-7 449.1 ms 2 quantization-undeclared
qwen/qwen3-embedding-4b 2560 $0.0200 $0.1807 ($0.1780–$0.1833) $1.73e-7 504.6 ms 1 single-endpoint-no-failover, quantization-undeclared, requires-documented-query-protocol
perplexity/pplx-embed-v1-4b 2560 $0.0300 $0.2693 ($0.2654–$0.2731) $2.30e-7 209.1 ms 1 single-endpoint-no-failover
voyageai/voyage-4 1024 $0.0600 $0.5385 ($0.5308–$0.5463) $4.60e-7 212.3 ms 1 single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared
mistralai/mistral-embed-2312 1024 $0.1000 $1.0910 ($1.0769–$1.1050) $1.07e-6 160.3 ms 3 quantization-undeclared
openai/text-embedding-ada-002 1536 $0.1000 $0.8746 ($0.8612–$0.8880) $7.00e-7 471.5 ms 1 single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared
voyageai/voyage-multimodal-3.5 1024 $0.1200 $1.0770 ($1.0616–$1.0925) $9.20e-7 269.7 ms 1 dimensions-silently-ignored, single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared
voyageai/voyage-4-large 1024 $0.1200 $1.0770 ($1.0616–$1.0925) $9.20e-7 240.8 ms 1 single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared
openai/text-embedding-3-large 3072 $0.1300 $1.1370 ($1.1195–$1.1544) $9.10e-7 513.8 ms 2 cross-endpoint-reordering, quantization-undeclared
google/gemini-embedding-001 3072 $0.1500 $1.4047 ($1.3785–$1.4309) $1.15e-6 361.6 ms 2 quantization-undeclared
mistralai/codestral-embed-2505 1536 $0.1500 $1.4203 ($1.4006–$1.4400) $1.50e-6 201.4 ms 3 quantization-undeclared
google/gemini-embedding-2 3072 $0.2000 $1.8729 ($1.8380–$1.9079) $1.53e-6 413.3 ms 4 quantization-undeclared
google/gemini-embedding-2-preview 3072 $0.2000 $1.8729 ($1.8380–$1.9079) $1.53e-6 335.3 ms 1 single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared

20 of 37 candidates were admitted; 17 were refused.

Refused, with the measurement that refused them#

Slug Refusals The measurement
liquid/lfm-2.5-embedding-350m:free trains-on-input truncation: refused-oversized-input; no endpoint accepts data_collection: deny
voyageai/voyage-code-4 identity-floor-failed, uncalibrated plain/nav-admin-repo-map first relevant at rank 6
nvidia/nemotron-3-embed-1b:free trains-on-input no endpoint accepts data_collection: deny
google/gemini-embedding-2:batch batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated serves only through the batch API; no cost-reported call
nvidia/llama-nemotron-embed-vl-1b-v2:free trains-on-input no endpoint accepts data_collection: deny
thenlper/gte-base silently-truncates, uncalibrated truncation: token-count-plateaus at 512 tokens
intfloat/e5-large-v2 silently-truncates, uncalibrated truncation: token-count-plateaus at 512 tokens
intfloat/e5-base-v2 silently-truncates, uncalibrated truncation: token-count-plateaus at 512 tokens
intfloat/multilingual-e5-large silently-truncates, uncalibrated truncation: token-count-plateaus at 512 tokens
sentence-transformers/paraphrase-minilm-l6-v2 silently-truncates, context-below-corpus-maximum, uncalibrated truncation: token-count-plateaus at 128 tokens
sentence-transformers/all-minilm-l12-v2 silently-truncates, identity-floor-failed, context-below-corpus-maximum, uncalibrated truncation: token-count-plateaus at 128 tokens; plain/id-admin-adr-0076 first relevant at rank 8
sentence-transformers/multi-qa-mpnet-base-dot-v1 silently-truncates, uncalibrated truncation: token-count-plateaus at 512 tokens
sentence-transformers/all-mpnet-base-v2 silently-truncates, uncalibrated truncation: token-count-plateaus at 384 tokens
sentence-transformers/all-minilm-l6-v2 silently-truncates, context-below-corpus-maximum, uncalibrated truncation: token-count-plateaus at 256 tokens
openai/text-embedding-ada-002:batch batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated serves only through the batch API; no cost-reported call
openai/text-embedding-3-large:batch batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated serves only through the batch API; no cost-reported call
openai/text-embedding-3-small:batch batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated serves only through the batch API; no cost-reported call

The identity floor#

Not a relevance measurement — task 3.5 owns those. This is the bar the cost rule implies: the rule binds "the cheapest model that CAN DO THE JOB", so a price ranking with no capability bar underneath it would rank models that cannot retrieve at all. Each candidate must put a page named by its own identifier above 200 chunks drawn from pages the label does not point at, using only the corpus-identifier labels task 3.1 settled without any ranker.

Several of these models are asymmetric by design and publish a query-side prefix. A model called without the one it documents is a model called wrong, so each is measured under every protocol its family documents AND under plain text as the control.

Slug Protocol Result
liquid/lfm-2.5-embedding-350m:free plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2
voyageai/voyage-code-4 none cleared plain/nav-admin-repo-map: FAIL @6; plain/id-admin-adr-0076: pass @1
voyageai/voyage-multimodal-3.5 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
voyageai/voyage-4-lite plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
voyageai/voyage-4 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
voyageai/voyage-4-large plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
nvidia/nemotron-3-embed-1b:free plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
google/gemini-embedding-2 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
google/gemini-embedding-2-preview plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
perplexity/pplx-embed-v1-4b plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
perplexity/pplx-embed-v1-0.6b plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3
nvidia/llama-nemotron-embed-vl-1b-v2:free plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
thenlper/gte-base plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
thenlper/gte-large plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
intfloat/e5-large-v2 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1
intfloat/e5-base-v2 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1
intfloat/multilingual-e5-large plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1
sentence-transformers/paraphrase-minilm-l6-v2 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3
sentence-transformers/all-minilm-l12-v2 none cleared plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @8
baai/bge-base-en-v1.5 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; bge-en-instruction/nav-admin-repo-map: pass @1; bge-en-instruction/id-admin-adr-0076: pass @1
sentence-transformers/multi-qa-mpnet-base-dot-v1 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @4
baai/bge-large-en-v1.5 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; bge-en-instruction/nav-admin-repo-map: pass @1; bge-en-instruction/id-admin-adr-0076: pass @1
baai/bge-m3 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
sentence-transformers/all-mpnet-base-v2 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2
sentence-transformers/all-minilm-l6-v2 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3
mistralai/mistral-embed-2312 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
google/gemini-embedding-001 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
openai/text-embedding-ada-002 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3
mistralai/codestral-embed-2505 plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1
openai/text-embedding-3-large plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2
openai/text-embedding-3-small plain plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2
qwen/qwen3-embedding-8b qwen3-instruct plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @7; qwen3-instruct/nav-admin-repo-map: pass @1; qwen3-instruct/id-admin-adr-0076: pass @2
qwen/qwen3-embedding-4b qwen3-instruct plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @16; qwen3-instruct/nav-admin-repo-map: pass @1; qwen3-instruct/id-admin-adr-0076: pass @1

The floor uses 2 corpus-identifier queries against 200 distractor chunks. Its protocols: plain (no prefix; the control every candidate is measured under); e5-query-passage (intfloat E5 family model cards: "query: " and "passage: " prefixes); bge-en-instruction (BAAI bge-*-en-v1.5 model cards: a query-side retrieval instruction); qwen3-instruct (Qwen3-Embedding model cards: an instruction-templated query side).

Provider data posture#

Posture is measured by asking the router for it. A request that routes had an endpoint meeting the filter; one that 404s had none, which the router could only know by checking every endpoint.

Slug Baseline Zero data retention Training denied
liquid/lfm-2.5-embedding-350m:free routed refused refused
voyageai/voyage-code-4 routed refused routed
voyageai/voyage-multimodal-3.5 routed refused routed
voyageai/voyage-4-lite routed refused routed
voyageai/voyage-4 routed refused routed
voyageai/voyage-4-large routed refused routed
nvidia/nemotron-3-embed-1b:free routed refused refused
google/gemini-embedding-2 routed routed routed
google/gemini-embedding-2-preview routed refused routed
perplexity/pplx-embed-v1-4b routed routed routed
perplexity/pplx-embed-v1-0.6b routed routed routed
nvidia/llama-nemotron-embed-vl-1b-v2:free routed refused refused
thenlper/gte-base routed routed routed
thenlper/gte-large routed routed routed
intfloat/e5-large-v2 routed routed routed
intfloat/e5-base-v2 routed routed routed
intfloat/multilingual-e5-large routed routed routed
sentence-transformers/paraphrase-minilm-l6-v2 routed routed routed
sentence-transformers/all-minilm-l12-v2 routed routed routed
baai/bge-base-en-v1.5 routed routed routed
sentence-transformers/multi-qa-mpnet-base-dot-v1 routed routed routed
baai/bge-large-en-v1.5 routed routed routed
baai/bge-m3 routed routed routed
sentence-transformers/all-mpnet-base-v2 routed routed routed
sentence-transformers/all-minilm-l6-v2 routed routed routed
mistralai/mistral-embed-2312 routed routed routed
google/gemini-embedding-001 routed routed routed
openai/text-embedding-ada-002 routed refused routed
mistralai/codestral-embed-2505 routed routed routed
openai/text-embedding-3-large routed routed routed
openai/text-embedding-3-small routed routed routed
qwen/qwen3-embedding-8b routed routed routed
qwen/qwen3-embedding-4b routed routed routed

Reranking, priced against what it would be added to#

Slug Served by Billing unit $/query Latency × the embedding query cost
qwen/qwen3-reranker-8b Fireworks unknown $0.0022886 567.2 ms 74628.26×
voyageai/rerank-2.5-lite VoyageAI by MongoDB unknown $0.00015334 389.8 ms 5000.22×
voyageai/rerank-2.5 VoyageAI by MongoDB unknown $0.00038335 358.5 ms 12500.54×
nvidia/llama-nemotron-rerank-vl-1b-v2:free Nvidia unknown $0 729.3 ms
cohere/rerank-4-pro Cohere search-unit $0.0025 521 ms 81521.74×
cohere/rerank-4-fast Cohere search-unit $0.002 366.5 ms 65217.39×
cohere/rerank-v3.5 Cohere search-unit $0.001 279.6 ms 32608.7×

deferred to task 3.4 — priced here and not adopted; the measured per-query cost is the reason it needs a quality argument before it is added, not a footnote.

One row bills nothing, and a zero there is a price, not a posture: every :free route in the embedding field was refused because no endpoint of it accepts data_collection: deny, and the refusal named "Free model training". A free reranker earns the same check before it is adopted.

The binding#

perplexity/pplx-embed-v1-0.6b is bound because it is the cheapest candidate that cleared every measured gate at 10,000 queries/month — $0.0040 per million tokens, 1024 dimensions, served by Perplexity. The runner-up is baai/bge-base-en-v1.5, $0.3352 per month dearer at that volume.

The binding is the slug, the endpoint set AND the query protocol plain — an index built under one convention and queried under another is two systems, not one.

An endpoint pin is still applied: this slug has a single endpoint, so no agreement between endpoints was measured and none is claimed; the pin names the endpoint that was.

Rollback#

  • Slug override: OSHUN_ASSISTANT_EMBEDDING_MODEL (registry supplies the default underneath it)
  • Endpoint override: OPENROUTER_PROVIDER_ONLY
  • Degraded mode: lexical-only — apps/oshun/bff/src/assistant/docs-search.ts
  • Rolling back invalidates stored vectors: true

What task 3.5 must measure#

Task 3.2 binds on price alone, and a price-only pin handed to 3.5 would leave it measuring one arm chosen by cost. The shortlist below is the admitted price/width frontier — for every admitted candidate NOT on it, some admitted candidate is at least as cheap and at least as wide. Dimensions are a capacity fact here, not a claim about retrieval.

  • perplexity/pplx-embed-v1-0.6b
  • qwen/qwen3-embedding-8b

What re-opens this binding#

  • task 3.5 measures relevance and a different candidate wins on a paired test
  • task 3.7 declines to promote dense retrieval at all, which retires this leg
  • the bound endpoint set changes, because vectors are only comparable within one endpoint
  • a re-priced route moves the chosen candidate out of first place at the recorded volume
  • the corpus chunker changes, because the calibrated tokens-per-character no longer applies
  • the query-side protocol changes, because an index built under one convention cannot be queried under another
  • the bound slug gains or loses an endpoint whose quantization differs from the measured one
  • the vendor re-points a versioned slug, because a vendor-versioned name is a pin only while the vendor keeps it one

Honest limits#

  • This binds on PRICE and POSTURE. No relevance measurement exists yet: task 3.5 owns Recall@k and nDCG over the full set with a pre-registered paired test, and task 3.7 owns promotion. A candidate that clears the identity floor has been shown able to retrieve, not shown to retrieve well.
  • The identity floor uses only the relevance set's corpus-identifier labels, because those are the ones that settle without any ranker. Four of the nine query classes still carry no labels at all (task 3.1), so no floor exists for conceptual, multi-hop, recency or contradictory queries.
  • Corpus token counts are a ratio estimator over disjoint sampled batches, not an exact count. The corpus CHARACTER count is exact; buying an exact token count would mean embedding all 36 million characters for every candidate.
  • Monthly figures are scenarios. This repository records no search_docs call rate, so the per-query and per-index prices are the measured facts and the monthly totals are arithmetic over a volume a reader chooses.
  • Latency is wall-clock from this host on this date, over a residential-grade path to one region. It is a floor for planning, not a service-level measurement.
  • Posture is measured as ROUTING behaviour — whether an endpoint meeting a filter exists — not as an audit of any provider's actual practice.
  • Nothing is indexed. Task 3.3 builds the index, and until it does the bound slug has no production consumer; search_docs remains lexical BM25F.
  • Prices are quoted for text as the corpus stores it. A binding whose query protocol adds a prefix costs a few tokens more per query and, where the protocol also prefixes documents, roughly two per cent more per index build; neither shifts the ranking between candidates at these magnitudes, and the protocol is recorded so the difference is recomputable.
  • Each price is the single usage.cost figure the route returned for one call, retained verbatim in the probe record beside its token count. This record therefore distinguishes no cached from uncached rate, and nothing here should be read as a claim that one exists or does not.
  • The identity floor is measured under the protocols this tool knows about. A model documenting a calling convention nobody here recognised is measured plain and would be under-measured exactly as the instruction-tuned models were before their protocols were added.