Date: 2026-06-16 Scope: the metis domain (libs/metis/**,
services/metis/**, apps/metis/**) — an adaptive-learning /
educational-content platform. Question (as set): analyze metis as it stands
today and identify missing features, especially around (1) SOTA knowledge &
correctness verification and (2) creation of rich, beautiful, compelling
media with autonomous agentic systems to teach concepts. Method: (a)
source-level audit of metis (real-vs-stub, file:line); (b) a fan-out /
adversarially-verified deep-research pass on the 2025–2026 SOTA for correctness
verification (23 verified findings) + a targeted pass on agentic teaching-media;
(c) cross-reference with the repo's existing platform reports
(V1_V9_AGENTIC_CONTENT_SOTA_ASSESSMENT_2026-06-14.md,
AGENTIC_CONTENT_QUALITY_ANALYSIS_2026-06-13.md).
0. Bottom line#
Metis is strong where most ed-tech is weak — real, end-to-end multimodal media orchestration (TTS with loudness normalization, video compositing, avatar + lip-sync, diagrams, presentations, multilingual narration) and real grounded governance (lesson-revalidation on source retraction, a rights/freshness filter, GraphRAG source inspection, a citation-required lesson schema). But it is weak exactly where education should be strongest: correctness verification is heuristic and single-pass, its named "fact-checker" does not actually check facts, and it generates no truly generative media (no math animation, no figure/image generation, no agentic author→critic→media loop).
Two findings dominate everything:
-
Education is the verifiable domain — and metis under-exploits that. The platform's own quality analysis nails the asymmetry: verifier-guided methods (best-of-N, process reward models, claim verification) deliver their big gains where a ground-truth verifier exists — math, code, formal reasoning; creative content has no oracle (
AGENTIC_CONTENT_QUALITY_ANALYSIS_2026-06-13.md §4). Educational content is precisely that domain. Yet metis verifies content with document-level heuristics and a single uncalibrated LLM judge, while the 2026 SOTA for verifiable content — atomic-claim decomposition, RAG faithfulness scoring, and process reward models — is absent. -
Metis does not ride the platform's verified SOTA content engine. A grep across
libs/metis,services/metis,apps/metisforcontent-quality-judge/content-release-gates/content-service/ grounding-gate / judge-panel returns nothing. The shared engine — a generate→judge-gate→refine loop with a calibrated, de-biased LLM-judge panel, best-of-N, a reward model with hacking detection, drift / champion-challenger, and a fail-loud canon-grounding gate that blocks ungrounded claims (V1_V9_AGENTIC_CONTENT_SOTA_ASSESSMENT_2026-06-14.md §2, 312+ tests) — sits in the same monorepo and metis reimplements a weaker subset of it.
Verdict by sub-area:
| Sub-area | Grade | One-line |
|---|---|---|
| Media orchestration/composition (TTS, audio, video, avatar, diagrams, decks) | A− | Genuinely strong, real, multimodal, multilingual. |
| Grounded governance (revalidation, rights, source inspection, cited-lesson schema) | B+ | Real and uncommon — a differentiator. |
| Correctness verification (factuality, faithfulness, math/STEM, judge quality) | D | Heuristic, single-pass; the "fact-checker" doesn't verify; math is syntax-only. |
| Generative teaching media (figures, math animation, video, sims) | D | Composits what it's given; generates ~none of it. |
| Agentic autonomy (plan→tool→observe→refine, author↔critic loops) | D | One-pass agents; no refine loop; shared engine unused. |
1. Current state of metis (verified)#
1.1 Correctness & verification — what's real, what isn't#
Real and good:
- Lesson revalidation on source retraction
(
research/src/credibility/lesson-revalidation.ts) — deterministic: counts credible supporting sources, applies a min-credibility threshold (≥0.6), disposes a lessonretained | flagged-for-review | quarantined. Wired to a veritas cascade contract. This is real, operational, and uncommon. - Rights/freshness governance (
knowledge-graph/.../governance.ts) —evaluateNode/Edgeblocksblocked/unknown-rights artifacts in learner contexts, requires source-backing, honors license/allowed-use. The GraphRAG benchmark suite (applications/graph-rag-benchmark.ts) inspects retrieved nodes for credibility, staleness, invalidation, rights, and source diversity. - Cited-lesson schema (
llm-client/.../content-generator) — the output schema requires ≥1 citation per lesson, with citation extraction from the body. - LLM-as-judge quality evaluator (
llm-client/.../quality-evaluator.ts) — a real single LLM call grading 5 pedagogical dimensions (completeness, readability, structure, relevance, Bloom alignment) with Zod validation + heuristic fallback. - Reading-level analysis (
quality/.../reading-level.ts) — real Flesch-Kincaid / Gunning Fog / Coleman-Liau / SMOG with complexity detection and linting rules. - Credibility scoring (
research/.../credibility.ts) — heuristic 5-factor with an exponential currency-decay by domain half-life.
The problem areas (verified by reading the code):
- The "fact-checking agent" does not check facts.
agents/.../fact-checking-agent.tsmakes no LLM call and no retrieval/search. It heuristically classifies a sentence's claim type, derives ascore, maps it to averdict(verified/disputed/false/ …) viadetermineVerdict(...), and returns hardcoded generic source categories (statistical: ['Government statistics databases', 'Published research studies']). It emits authoritative-sounding verdicts it never grounded in evidence — a correctness / credibility hazard, and exactly the result-faking pattern the repo's quality rules forbid. - Math validation is syntax-only.
quality/.../math-validator.tschecks balanced brackets / well-formed LaTeX (tagged['math','syntax']). There is no symbolic / CAS / step / correctness checking — in the one place education has a true oracle. - The quality judge is single-pass and uncalibrated. One
isis.complete()call; no panel, no position-bias mitigation, no calibration against human gold — the very failure modes the 2026 judge literature flags (see §2.1). - Citations are required but never verified to support their claims.
quality/.../citation-checker.tsvalidates citation format (APA/MLA), not whether a cited source actually entails the sentence. Citation existence ≠ citation sufficiency. - Search/embeddings/corpus/citations modules in
research/are largely types/factories (interfaces + constructors), not implementations.
1.2 Teaching media — what's real, what's absent#
Real and strong (this is metis's best work):
- Audio/TTS (
multimedia/src/audio/*): real TTS orchestration (SSML, provider abstraction incl. ElevenLabs/Chatterbox/local, emotion-driven prosody, ffmpeg format conversion), ITU-R BS.1770 LUFS loudness normalization, silence/fingerprint/dup detection, narration planning, phonetic customization, multilingual narration. - Video (
multimedia/src/video/*): frame-level compositing, transitions, captions, GPU worker pool, avatar generation + phoneme→viseme lip-sync, chroma-key/PiP/audio-mix compositor. - Diagrams (
multimedia/src/diagram/*+agents/diagram-agent.ts): SVG rendering with theming, progressive-reveal animation, interactive tooltips, Mermaid/SVG/JSON output. - Presentations & lectures (
multimedia/src/presentation/*,lecture-generation/*,avatar-teacher/*): deck assembly, speaker notes, multi-format export (PDF/HTML/MP4), end-to-end lecture orchestration (slides + narration + diagrams + transcripts), avatar teacher synthesis. The Python service addslecture_generation_serviceand amedia_evaluation_service(audio/video/caption/sync/embodiment quality).
Absent vs SOTA:
- No generative figure/diagram-from-meaning — diagrams are rendered from
caller-specified specs; there is no text→image generation for concept
illustrations (would route via
@isis/Yemaya — currently only stub adapters inintegrations/yemaya-integration.ts). - No programmatic math/CS animation (Manim-class). The single highest-value ed-media gap.
- No generative (pixel) text-to-video (Sora/Veo-class) — only compositing of supplied assets.
- No interactive simulations/widgets (PhET/Desmos-class).
- No multi-agent author→critic→media refinement loop. Agents are one-pass generators; there is no plan→generate→render→critique→repair cycle.
- No accessibility automation — transcripts/captions exist, but no auto alt-text for diagrams/figures, no audio descriptions, no sign-language track.
2. The 2026 SOTA yardstick (cited)#
2.1 Correctness & knowledge verification#
- Atomic-claim decomposition is the factuality paradigm. Break a generation into atomic facts, verify each against a knowledge source, aggregate. FActScore = % of atomic facts supported (automated retrieval+LLM estimator <2% error vs humans on biographies; higher cross-domain) [arXiv:2305.14251]. SAFE verifies each atomic fact via multi-step Google-Search reasoning and scores F1@K (precision of supported facts vs recall to a target length K) [arXiv:2403.18802]. FaStfact (EMNLP 2025) cuts extraction from O(N) to O(N/w) via chunk-level extraction + confidence pre-verification, with the tightest human alignment among evaluated pipelines [arXiv:2510.12839].
- LLM-as-judge is unreliable unless engineered. Systematic position bias varying by judge/task [arXiv:2406.07791] and an "agreeableness" bias — judges confirm correct outputs at TPR >96% but reject invalid ones at TNR <25% [arXiv:2510.11822]. SOTA mitigations: Cascaded Selective Evaluation (conformal, distribution-free guarantee of human-agreement by abstaining/escalating; cheap models exceed 80% agreement where GPT-4 gets 77.8%) [arXiv:2407.18370]; minority-veto and regression bias-correction ensembles (max error 2.8%→1.2%, beating a 14-LLM majority) [arXiv:2510.11822]; IRT/Graded-Response-Model reliability diagnostics — no judge is uniformly reliable across criteria, validate per task [arXiv:2602.00521]; linear probes on judge hidden states for calibrated confidence at ~10× less compute than self-consistency [arXiv:2512.22245].
- RAG faithfulness/groundedness has dedicated scorers. A fine-tuned DeBERTa-v3 NLI judge reaches hallucination-AUROC 0.64–0.86, beating prompted GPT-3.5 (0.51–0.65), RAGAS, TruLens; the TRACe framework (RAGBench) defines context-relevance, utilization, completeness, and adherence (= faithfulness/groundedness/attribution: are all parts of the response grounded in context?) [arXiv:2407.11005].
- Citations: retrieval is the dominant driver; placement is a trade-off. Zero-shot→RAG is the biggest citation-correctness gain (FEVER 27%→77%); post-hoc citation maximizes coverage, generation-time maximizes precision [NeurIPS 2025, arXiv:2509.21557].
- STEM correctness uses process reward models. ThinkPRM — a generative long-CoT verifier fine-tuned on ~1% of PRM800K labels — beats discriminative PRMs trained on the full set [arXiv:2504.16828]. PRM800K is the canonical process-supervision dataset.
2.2 Autonomous agentic teaching-media generation#
- Programmatic, agent-authored math/CS animation is the education-native SOTA. TheoremExplainAgent generates long-form (>5 min) theorem-explanation Manim videos; on TheoremExplainBench (240 theorems, 5 metrics) an o3-mini agent hits 93.8% success / 0.77 — and crucially finds that multimodal video explanations expose reasoning gaps that text-only explanations hide [arXiv:2502.19400].
- Multi-agent author→critic→media is the pipeline pattern. Code2Video uses a Planner → Coder → Critic trio to generate educational videos via executable code (not pixel synthesis), with a VLM critic refining layout/clarity; +40% over direct code generation, comparable to human tutorials, evaluated on MMMC with the TeachQuiz metric (does a VLM learn from the video?) [arXiv:2510.01174].
- Generative video is production-ready in 2026. Sora 2 (physics realism, native audio), Veo 3.1 (4K, native audio), Kling 3, Seedance 2.0 — single-API, multi-shot, native audio.
- Talking-head avatars are omnimodal. Hedra Character-3 (joint image+text+audio), HeyGen Avatar 5 (15-s training, 1080p, 175+ languages), ByteDance OmniHuman v1.5 (film-grade); joint audio-video models beat separate-step lip-sync (EvalTalker, arXiv:2512.01340). (Metis already has avatar+lip-sync orchestration — the gap is the generative model behind it, not the pipeline.)
- Baseline vs SOTA distinction: baseline = template/compositing of supplied assets; SOTA = an agent that plans a lesson, generates the figures/animation/voiceover, renders, and a critic loops until a learning-outcome proxy (e.g. TeachQuiz) passes — with accessibility (alt-text, captions) produced as a first-class output.
3. Gap analysis — Area 1: correctness verification (prioritized)#
| # | Gap | SOTA reference | Severity |
|---|---|---|---|
| C1 | Replace the non-verifying fact-checker with real claim-decomposition verification. Decompose lessons into atomic claims; verify each against the metis knowledge graph + retrieval; aggregate to a FActScore/F1@K-style groundedness score; surface unsupported claims for review/regeneration. | FActScore [2305.14251], SAFE/F1@K [2403.18802], FaStfact [2510.12839] | P0 — credibility-critical (today it fabricates verdicts) |
| C2 | RAG faithfulness / citation-sufficiency scoring. Score whether each cited source actually entails the sentence it backs (TRACe "adherence"); a fine-tuned NLI (DeBERTa) or calibrated judge, not format-checking. Turns "≥1 citation required" into "citations actually support the claims." | TRACe/RAGBench [2407.11005]; citation placement [2509.21557] | P0 |
| C3 | Real STEM correctness, not syntax. Add symbolic/CAS checking (e.g. SymPy/Mathematica-class) for math, execution-based checking for code, and a process reward model (ThinkPRM-style) to verify worked-solution steps — exploiting education's ground-truth oracle. | ThinkPRM [2504.16828], PRM800K | P0 (STEM) |
| C4 | Upgrade the single judge to a calibrated, de-biased panel — or adopt the shared engine's. Position-bias mitigation, ≥3 judges with disagreement→HITL, calibration vs human gold, IRT reliability per criterion, linear-probe confidence. The platform already has a verified one; reuse it. | [2406.07791, 2510.11822, 2407.18370, 2602.00521, 2512.22245]; shared content-quality-judge |
P0 |
| C5 | Contradiction / cross-source conflict detection ("does source A refute B?") and claim-to-source span linking (fine-grained, not document-level). | TRACe; NLI entailment/contradiction | P1 |
| C6 | Uncertainty calibration — emit confidence intervals / calibrated scores, not point estimates; route low-confidence content to HITL (active learning). | linear-probe calibration [2512.22245] | P1 |
| C7 | Verifier-guided generation loop — best-of-N + the verifier(s) above as the selector, then self-refine on failed claims. Education's oracle makes this far more effective than in creative domains. | quality-analysis §4; ThinkPRM | P1 |
| C8 | Pedagogical/curriculum correctness — age-appropriateness, prerequisite-consistency (tie to the knowledge graph), misconception checks (a misconception-graph already exists — wire it into verification). |
open question in research; reuse internal KG | P2 |
4. Gap analysis — Area 2: rich agentic teaching media (prioritized)#
| # | Gap | SOTA reference | Severity |
|---|---|---|---|
| M1 | Agentic programmatic math/CS animation (Manim). metis has a Python service and a diagram-agent — a TheoremExplainAgent/Code2Video-style Planner→Coder(Manim)→Critic loop is a natural, high-impact addition; renders to the existing video pipeline. | TheoremExplainAgent [2502.19400], Code2Video [2510.01174] | P0 — highest-value media gap |
| M2 | Multi-agent author→critic→media refinement loop with a learning-outcome proxy (TeachQuiz-style: can a model learn from the artifact?) gating release — replaces one-pass agents. | Code2Video [2510.01174] | P0 |
| M3 | Generative figures/diagrams from meaning (text→image for concept art, generative charts from data) — wire the Yemaya/Isis adapters (currently stubs) to real image models, gated by the verification stack. | 2026 image models; integrations stubs | P1 |
| M4 | Generative explainer video (Sora/Veo/Kling-class) for segments where compositing/animation is insufficient — via the Isis/Yemaya provider seam, not in-domain. | Sora 2 / Veo 3.1 / Kling 3 | P1 |
| M5 | Accessibility automation as a first-class output — auto alt-text for every diagram/figure (VLM-generated, then verified), audio descriptions, caption QA (the media_evaluation_service can gate it). | WCAG; VLM alt-text | P1 |
| M6 | Interactive simulations/widgets (PhET/Desmos-class) generated/parameterized per concept. | PhET/Desmos | P2 |
| M7 | Learned, adaptive narration pacing (from learner attention/comprehension) rather than templated timing. | — | P2 |
5. The cross-cutting gap (ties both areas together)#
Both areas share one root cause: metis runs one-pass agents and its own
quality subset instead of the platform's verified, closed-loop SOTA engine.
The shared generate → judge-gate → (regenerate | refine) → re-gate loop, the
calibrated judge panel, the reward model with hacking detection,
drift/champion-challenger, and the canon-grounding gate that blocks ungrounded
claims are real and tested
(V1_V9_AGENTIC_CONTENT_SOTA_ASSESSMENT_2026-06-14.md §2) — and unused by
metis. The same loop is exactly what both a correctness-verifier (C1–C4, C7) and
an author→critic→media pipeline (M1–M2) need. The single highest-leverage move
is to put metis content and media generation on that loop, with
education-specific verifiers (claim/RAG/PRM/CAS) plugged in as the gate —
turning "generate then a human checks" into "generate, machine-verify against
ground truth, refine, then a human handles only low-confidence cases."
The platform quality report's standing warnings apply directly to metis: don't scale generation before the quality/verification layer exists; and the human reviewers' accept/ reject decisions should train the verifier (active learning), not be warehoused.
6. Prioritized roadmap#
- P0 (correctness-critical, do first): C1 (claim-decomposition verifier replacing the fake fact-checker) · C2 (RAG faithfulness/citation-sufficiency) · C3 (real math/code/step correctness) · C4 (calibrated judge panel — reuse the shared engine) · M1 (agentic Manim animation) · M2 (author→critic→media loop). Sequence: stand the verifiers up as a gate, put generation on the shared loop, then add the Manim media agent behind the same gate.
- P1: C5–C7 (contradiction/span linking, calibration, verifier-guided best-of-N) · M3–M5 (generative figures, explainer video via Isis/Yemaya, accessibility automation).
- P2: C8 (pedagogical/curriculum correctness via the KG + misconception-graph) · M6–M7 (interactive sims, adaptive pacing).
7. What metis genuinely gets right (so this isn't only critique)#
Real, end-to-end multimodal media orchestration (TTS + LUFS audio mastering + video compositing + avatar/lip-sync + diagrams + decks + multilingual lectures) is uncommon and production-shaped. Grounded governance — lesson revalidation on source retraction, a context-aware rights/freshness filter, GraphRAG source inspection, and a citation-required lesson schema — is ahead of typical ed-tech and is the right foundation to bolt real verification onto. The gap is not the plumbing; it's the verification intelligence and the generative-media intelligence that should ride on top of it.
Appendix — sources & caveats#
Area 1 (3-vote verified): FActScore arXiv:2305.14251 · SAFE/F1@K
arXiv:2403.18802 · FaStfact arXiv:2510.12839 · position bias arXiv:2406.07791 ·
agreeableness bias + ensembles arXiv:2510.11822 · Cascaded Selective Evaluation
arXiv:2407.18370 · IRT judge diagnostics arXiv:2602.00521 · linear-probe
calibration arXiv:2512.22245 · RAGBench/TRACe arXiv:2407.11005 · citation
placement arXiv:2509.21557 · ThinkPRM arXiv:2504.16828. Area 2 (targeted pass;
the automated deep-research run failed to verify Area-2 claims — these are from
primary papers + 2026 model comparisons, treat model-ranking deltas as
directional): TheoremExplainAgent arXiv:2502.19400 · Code2Video
arXiv:2510.01174 · talking- head eval arXiv:2512.01340; Sora 2 / Veo 3.1 / Kling
3 / Hedra Character-3 / HeyGen Avatar 5 / OmniHuman v1.5 per 2026
vendor/benchmark comparisons. Internal evidence:
libs/metis/agents/src/agents/fact-checking-agent.ts (heuristic verdicts, no
retrieval/LLM, hardcoded source categories), quality/.../math-validator.ts
(syntax-only), llm-client/.../quality-evaluator.ts (single isis.complete),
research/.../credibility/lesson-revalidation.ts,
knowledge-graph/.../governance.ts,
multimedia/src/{audio,video,diagram,presentation}/*,
integrations/src/yemaya-integration.ts (stub adapters); grep confirms metis
does not import content-quality-judge / content-release-gates /
content-service. Caveats: several SOTA headline numbers are authors'
self-reports on their own benchmarks (FaStfact, Code2Video, Cascaded Selective
Evaluation); agreeableness-bias / ensemble magnitudes derive from one narrow
task; how these verifiers perform on pedagogical correctness (vs
biography/code domains) is an open question; the Area-2 automated verification
did not complete, so its system rankings are directional.