Mawu · Guides & deep dives

V7 IP Content Matching — A Content-ID-Class Plan Beyond Hash Lists

FiveM's content economy floats on ripped AAA assets: vehicles, clothing, and props converted from GTA V, Forza, CSR, and other commercial games, traded and sold at scale.

8sections9 minread3tables

On this page

Status: V7 trust-and-safety / rights design. Source: fills the planning gap identified in V1_V7_PLAN_SET_AUDIT_2026-06-12.md §6.2 (hash lists alone cannot police the ripped-asset economy). Grounded in V7/V7_features.md (features§"Commerce and Rights" — the V5 known-infringing-asset-hash list, features§"Sekhmet — Malware Scanning and Supply-Chain Integrity", features§"Mawu Studio — Creation Tools" asset pipeline) and V7/V7_ARCHITECTURE.md (arch§"Sekhmet — Trust, Safety, and Child-Protection Architecture" — the render-3D-to-2D precedent, arch§"Dependency Resolution and Content-Addressed Storage"). Owner: Sekhmet T&S lead (pipeline), rights/legal lead (portal + notices), Themis (disputes). Date: 2026-06-12.

1. Why Hash Lists Are Not Enough#

FiveM's content economy floats on ripped AAA assets: vehicles, clothing, and props converted from GTA V, Forza, CSR, and other commercial games, traded and sold at scale. Exact-hash blocklists (the V5 mechanism carried in features§"Commerce and Rights") stop byte-identical re-uploads only; the ripped-asset pipeline routinely defeats them with re-exports, retopology, decimation, scale/UV changes, and texture recompression — every one of which changes every byte while preserving the infringement. A platform that wants the clean-market position promised to original creators (V7/docs/community-migration.md §2.3) needs perceptual matching — the Content ID model: a reference corpus from rightsholders, similarity matching at ingest, confidence-tiered actions, and a dispute flow.

Scope: this system targets commercial game/film/music IP infringement. It is separate from CSAM hashing (arch§"Sekhmet", csam-hash gate) — different databases, different legal regime, different teams — but deliberately reuses the same architectural slot: a platform-central scan at ingest, before any client can fetch (features§"Sekhmet").

2. Perceptual Matching for 3D Assets#

Every uploaded mesh asset runs a three-channel fingerprint at Sekhmet ingest; channels are independent so evasion must beat all of them simultaneously.

2.1 Geometry channel — normalized shape descriptors#

Robust to retopology, decimation, scale, rotation, and translation:

  1. Canonical normalization: center to centroid; scale to unit bounding sphere; PCA pose alignment (principal axes, sign-disambiguated by third-moment) so descriptor extraction is orientation-invariant.
  2. Descriptors (computed on the normalized mesh, stored as the asset's geometry fingerprint):
    • D2 shape distribution (Osada-class): histogram of distances between ~1M random surface point pairs, 256 bins — cheap, retopo-invariant, coarse.
    • Spherical-harmonics descriptor: rotation-invariant SH power spectrum of the voxelized shape function (64³ grid, bands 0–16) — robust mid-grain shape signature.
    • Multi-view descriptor: the asset rendered from the 26 canonical viewpoints (vertices+face-centers of a rhombicuboctahedron-style rig) — deliberately the same render-canonical-views machinery Sekhmet already builds for 3D CSAM screening (arch§"Sekhmet": "render canonical 2D views and hash those"), so this channel shares infrastructure rather than duplicating it. Each view feeds the §2.3 embedding and a PDQ-class perceptual hash.
  3. Matching: D2/SH distance gates a cheap pre-filter; candidates within the pre-filter threshold go to the embedding ANN index (§2.3) for the scored match.

2.2 Texture channel — perceptual hashes#

Every texture/material map (albedo first; normal/roughness as secondary corroboration) gets a PDQ-class perceptual hash plus a 64-bit pHash; matching is Hamming-distance against the reference corpus's texture hashes. Texture matches alone are corroborating evidence (textures are the most-reused and also most-legitimately-similar channel); geometry+texture agreement is the high-confidence combination. Livery/decal regions are masked out of the hash where the material graph marks them (custom liveries on an infringing vehicle body must not launder the body; the body still matches).

2.3 Embedding channel — learned similarity vs. a reference corpus#

A PointNet++/DGCNN-class point-cloud encoder plus a multi-view CNN encoder (trained with contrastive pairs: original vs. decimated/retopologized/ re-exported variants of licensed training meshes) produces a 512-d embedding per asset. Reference-corpus embeddings live in an ANN index (HNSW; the pgvector/Qdrant infrastructure already in the monorepo dev stack). Cosine similarity against the corpus is the primary match score. Planning assumption adopted 2026-06-12: encoder training data is built from (a) rightsholder reference uploads (§4), (b) licensed/owned asset libraries, (c) public-domain and CC0 corpora — never from scraping commercial games ourselves; the reference corpus for matching is rightsholder-supplied or platform-licensed, which is what makes enforcement defensible.

2.4 Audio fingerprinting#

Two surfaces:

  • Uploaded audio assets (music in resources, radio-station packs — the classic FiveM infringement): Chromaprint-class constellation fingerprinting at ingest against a licensed commercial-music reference DB (Audible Magic/Pex-class vendor integration rather than building the music corpus in-house; planning assumption adopted 2026-06-12 — vendor selection is a procurement task, budget line in §7).
  • In-realm streamed music (realm "radio" resources pulling external streams): the resource capability model already gates network access (features§"The Six Sandbox Tiers" — Scripted has no network); realms wanting licensed-music streaming use the platform media capability, which fingerprints server-side at the media service, sampling 1 window per 30 s of stream. DMCA exposure for live realm audio is handled the way live platforms do: fingerprint, mute-on-match for the affected channel, log for the rightsholder portal.

3. Match-Confidence Thresholds and Actions#

Calibrated on a labeled validation set (known rips + hard-negative originals); numbers below are the launch calibration targets — planning assumption adopted 2026-06-12, re-calibrated quarterly against the false-positive metric.

Confidence Action at ingest
High — exact/known hash, OR embedding cosine ≥0.92 with geometry pre-filter agreement, OR geometry+texture dual-channel match Block at ingest. Artifact never becomes fetchable (the Sekhmet posture: scanned before any other machine can fetch). Uploader notified with the matched reference (rightsholder name, not internal data) + counter-notice path (§5).
Medium — single-channel match, cosine 0.80–0.92, or audio partial match Flag for human review in the Kuanyin rights queue; artifact held in InModeration (the existing publish lifecycle, features§"The In-Game Realm and Mod Browser"); reviewer SLA 72h; reviewer outcome trains the calibration set.
Low — cosine <0.80, no corroboration Pass; fingerprint + score logged. Logged low scores against the same reference family feed repeat-pattern review (an uploader with twenty 0.75s against the same vehicle line gets a human look).

Repeat-infringer policy (DMCA 512(i) requirement): strikes at confirmed infringement; 3 confirmed strikes in 12 months = upload privileges revoked, monetization clawback per the Aje reserve mechanism; appeal via Themis.

4. The Rightsholder Portal#

A web surface (extends the Mawu Operator Console family, apps/oshun/admin/src/app/mawu/) for verified rightsholders:

  1. Onboarding/verification: corporate identity verification before any reference access or takedown power (Content ID's lesson: portal abuse is its own abuse class).
  2. Reference uploads: meshes/textures/audio they own, fingerprinted into the reference corpus (fingerprints retained; reference media held encrypted, access-logged, used only for matching and reviewer side-by-side). Bulk API for catalog-scale rightsholders.
  3. Match review dashboard: matches against their references with confidence, status (blocked/under review/counter-noticed), and the reviewer-facing evidence (side-by-side canonical views, overlap heatmaps).
  4. Action requests: confirm block, request takedown of a Medium match, or whitelist (e.g., an official partnership or a licensed creator program — license records attach to the reference so licensed creators stop matching).
  5. Standard DMCA intake remains available for non-onboarded rightsholders (features§"Commerce and Rights": IP-claim and DMCA path routes through Themis); the portal is the scaled path, not the only path.

5. Counter-Notice and Fair-Use Handling#

Modeled on DMCA 512(g) with the platform's Themis adjudication on top:

  1. Uploader receives the block/takedown with the specific match basis and may file a counter-notice (identity, good-faith statement, consent to jurisdiction — templated in-product).
  2. Counter-noticed items go to human review with escalation to Themis; automated matching alone never decides a contested case. Reviewers apply a documented rubric: independent-creation evidence (working files, version history — Mawu Studio provenance metadata, features§"Mawu Studio", is strong evidence here), license claims, and the transformation/parody/commentary considerations of fair use. The platform does not pretend to adjudicate fair use as a court would; the rubric targets "plausible enough to restore and route to the statutory process."
  3. Restore timeline: if the rightsholder does not initiate legal action within the statutory 10–14 business-day window after a valid counter-notice, the content is restored.
  4. Provenance head start: assets authored in Mawu Studio carry creation provenance (features§"Mawu Studio — Creation Tools"); imported assets don't. Imported assets therefore face the matching bar with less exculpatory metadata — documented in creator guidance so original creators know keeping working files matters.
  5. EU note: for EU users this flow doubles as the DSA Art. 17/20 statement-of-reasons + internal-complaint mechanism (the platform already carries DSA appeal duties, features§"Content Moderation and Minor Protection").

6. Phased Rollout#

Phase Timing (relative to creator alpha) Contents
P0 — Launch Creator alpha day 1 Known-infringing hash list (V5 mechanism) + exact content-address dedup (free, via arch§"Dependency Resolution") + DMCA intake via Themis + repeat-infringer policy
P1 — Texture + audio perceptual +2 months PDQ/pHash texture matching; audio vendor fingerprinting at ingest; rightsholder portal beta with 3–5 anchor rightsholders (planning assumption adopted 2026-06-12)
P2 — 3D perceptual +6 months (gates economy launch for paid 3D listings) Geometry descriptors + multi-view + embedding ANN at ingest; thresholds per §3; full portal GA; back-scan of the existing catalog (§6.1)
P3 — Continuous ongoing Quarterly threshold recalibration; encoder retraining as evasion evolves; reference-corpus growth targets; in-realm stream sampling

6.1 Back-scan#

When P2 lands, the entire existing catalog is re-scanned (content-addressing makes this enumerable and incremental). Matches follow §3 but with a 14-day notice-before-delisting for already-published items, because retroactive instant blocks on realms' pinned lock files would break live communities — quarantine honors the no-silent-deletion-of-pinned-versions rule (features§"Realm Lifecycle and Registry") by flagging and de-monetizing first, hard-blocking after the notice window.

7. Scanning Cost Estimate per Upload#

Planning assumptions adopted 2026-06-12; validated against the P2 prototype before budget commit. Per 3D asset at ingest:

  • 26 canonical-view renders at 512²: ~1.5 s GPU (amortized batch)
  • Point sampling + D2 + SH descriptors: ~1.0 s CPU
  • Embedding forward passes (point-cloud + multi-view): ~0.7 s GPU
  • ANN query + texture hashes: ~0.3 s
  • Total ≈ 3–4 s mixed GPU/CPU ≈ $0.002–0.006 per asset at on-demand cloud GPU rates (T4/L4-class at $0.35–0.80/hr, batched), plus storage of ~6 KB of fingerprints per asset. At 10,000 asset uploads/day (an aggressive year-1 ceiling), worst case ≈ $60/day ≈ $1.8k/month of scan compute — negligible against the legal exposure it controls. Audio vendor fingerprinting is priced per-lookup (vendor-dependent, est. $0.001–0.01/lookup; procurement to confirm).
  • Sekhmet scan-queue latency budget already requires publish-to-available to be predictable (features§"Performance Budgets"); the IP pass shares the queue and must fit the same bound — target p95 ≤10 min added latency at P2.

8. Metrics and the Eval Gate#

In the V7/ADVERSARIAL_EVAL_GATES.md style — add row when the Sekhmet test rig grows the corpus (owner: Sekhmet lead):

Gate Owner Threshold Adversarial proof
ip-match Sekhmet ≥95% recall on the seeded rip corpus (exact + decimated + retopo + rescaled + retextured variants); ≤1% false-positive on the hard-negative original corpus; blocked items never fetchable Corpus includes adversarial variants (vertex jitter, UV scramble, mirrored geometry, livery overlays) that must still match; an original-but-similar hard-negative set must pass

Operational metrics: match volume by confidence band, reviewer overturn rate (target <10% of High blocks overturned — else thresholds are wrong), counter-notice rate and restore rate, time-to-takedown for portal requests (target ≤72h), repeat-infringer interdiction rate, and rightsholder corpus coverage (number of onboarded rightsholders / reference assets).