Overview#
Nisaba is the Ancient Text Analysis and Cross-Tradition Scholarly Research domain. Its mission is to give scholars a single, coherent environment for working with pre-modern texts across every tradition — cuneiform tablets, Sanskrit manuscripts, Hebrew scrolls, Quranic codices, Greek papyri, Chinese classics, and dozens of other writing systems. The real-world problem Nisaba solves is fragmentation: primary sources, lexica, digital corpus platforms, annotation tools, and critical-edition software all live in separate silos with incompatible data models. Nisaba replaces that patchwork with a unified, typed, interoperable library stack.
The domain lives entirely in libs/nisaba/ within the Oshun Nx monorepo — there
is no apps/nisaba or services/nisaba. It is organized as 22 independently
publishable packages following the @nisaba/<name> convention. The platform
serves scholars across biblical studies, classics, Assyriology, Indology,
Buddhist textual studies, Quranic manuscript traditions, and any field engaged
with pre-modern written heritage.
Because Nisaba ships as pure libraries (no app, no service), the packages are consumed by other Oshun domains (Tara, Arete, Veritas, Nyx) and by external scholarly tooling through the generated V1 REST contract. This design keeps the scholarly logic separate from UI and service concerns and allows each library to be versioned and published independently.
Library Organization#
The 22 packages form a layered stack ranging from low-level type vocabulary and Unicode utilities all the way up to client SDKs, mobile surfaces, and cross-domain bridges.
libs/nisaba/
├── schemas/ # @nisaba/schemas — Zod validation schemas for domain entities
├── core/ # @nisaba/core — Types, constants, utilities, error hierarchy
├── languages/ # @nisaba/languages — Ancient writing system handlers
├── annotations/ # @nisaba/annotations — W3C WADM standoff annotation engine
├── assistant/ # @nisaba/assistant — AI translation commentary and variant evaluation
├── canon/ # @nisaba/canon — Canonical text reference systems (all traditions)
├── client/ # @nisaba/client — Scholarly SDK and React hooks
├── api-client/ # @nisaba/api-client — Generated typed client for the V1 contract API
├── comparative/ # @nisaba/comparative — Cross-tradition concept mapping and parallels
├── corpora/ # @nisaba/corpora — Corpus connectors and ingestion pipeline
├── criticism/ # @nisaba/criticism — Textual criticism: collation, stemma, apparatus
├── cross-domain/ # @nisaba/cross-domain — Cross-domain bridge (Tara, Arete, Veritas, Nyx)
├── database/ # @nisaba/database — Prisma client and schema
├── editions/ # @nisaba/editions — Critical edition project management
├── geotemporal/ # @nisaba/geotemporal — Geographic timeline and provenance mapping
├── mobile/ # @nisaba/mobile — Mobile study surface definitions
├── paleography/ # @nisaba/paleography — ML-assisted manuscript analysis
├── philology/ # @nisaba/philology — Linguistic analysis: etymology, semantics, stylometry
├── standards/ # @nisaba/standards — TEI, Linked Data, encoding standard compliance
├── study-plans/ # @nisaba/study-plans — Study plan model and forecasting
├── translations/ # @nisaba/translations — Translation comparison, alignment, divergence
└── workspace/ # @nisaba/workspace — Research environment: split-pane layout, search
Dependency Layering#
Understanding the dependency graph matters when deciding where new code belongs.
@nisaba/core is the single foundation: it carries the shared type vocabulary
and the error hierarchy, and its only external dependency is @oshun/errors.
The remaining Nisaba packages keep a deliberately shallow dependency graph —
most either depend only on @nisaba/core or declare no intra-domain
dependencies at all, which prevents circular imports and allows packages to be
built and published independently.
The table below lists the actual intra-domain dependency edges, confirmed from
package.json and source imports:
@nisaba/core ← external: @oshun/errors (no Nisaba deps)
@nisaba/schemas ← external: zod; Nisaba: @nisaba/core
@nisaba/languages ← @nisaba/core
@nisaba/assistant ← @nisaba/core
@nisaba/canon ← @nisaba/core
@nisaba/criticism ← @nisaba/core
@nisaba/geotemporal ← @nisaba/core
@nisaba/paleography ← @nisaba/core
@nisaba/philology ← @nisaba/core
@nisaba/translations ← @nisaba/core
@nisaba/cross-domain ← @nisaba/study-plans
@nisaba/annotations ← (no Nisaba deps)
@nisaba/comparative ← (no Nisaba deps)
@nisaba/corpora ← (no Nisaba deps)
@nisaba/editions ← (no Nisaba deps)
@nisaba/standards ← (no Nisaba deps)
@nisaba/workspace ← (no Nisaba deps)
@nisaba/database ← (no Nisaba deps)
@nisaba/client ← (no Nisaba deps; zero-dependency SDK)
@nisaba/api-client ← (no Nisaba deps; generated from the OpenAPI types)
@nisaba/mobile ← (no Nisaba deps)
@nisaba/study-plans ← (no Nisaba deps)
Packages listed as "no Nisaba deps" are presently self-contained — their
package.json declares no @nisaba/* dependency and their source does not
import one. This layering is not yet enforced by an ESLint boundary rule.
Core Package Design#
@nisaba/core#
@nisaba/core is the type vocabulary shared across the entire domain. It
provides the enums, interfaces, utility functions, and error classes that every
other Nisaba package builds on. New code that needs a shared type should add it
here rather than duplicating definitions across packages. Its internal
structure:
core/src/
├── types/
│ ├── writing-systems.ts — WritingSystem enum (ISO 15924), ScriptDirection, WritingSystemMetadata
│ ├── morphology.ts — PartOfSpeech, VerbStem, MorphologicalParse, grammatical-feature unions
│ ├── reference.ts — ReferenceSystemType, CanonicalRef, CanonicalRefRange, CanonicalReferenceSystem
│ ├── traditions.ts — Tradition enum, TraditionMetadata
│ ├── manuscript.ts — ManuscriptMaterialType, VariantClassificationType
│ ├── apparatus.ts — ApparatusFormatType, VariantClassification, ApparatusEntry, ApparatusReading
│ ├── annotations.ts — AnnotationTypeValue, AnnotationMotivation, W3C selectors
│ ├── lexicon.ts — LexiconEntry, LexicalSense, LexicalCrossRef
│ ├── transliteration.ts — TransliterationScheme enum, TransliterationConfig
│ └── comparative.ts — SimilarityTaxonomy, EvidenceGradeType, ConsensusLevelType, ComparativeEvidence
├── constants/
│ ├── script-metadata.ts — SCRIPT_METADATA registry per WritingSystem
│ ├── unicode-ranges.ts — UNICODE_RANGES per writing system
│ └── tradition-metadata.ts — TRADITION_METADATA registry
├── utils/
│ ├── unicode.ts — NFC normalization, range checks, script detection
│ ├── offsets.ts — Byte/character offset conversions
│ ├── dates.ts — Approximate BCE/CE date parsing
│ ├── confidence.ts — Confidence scoring utilities
│ ├── citations.ts — Citation formatting
│ └── references.ts — nisaba:// URI parsing and construction
└── errors.ts — NISABA_ERROR_CODES registry + NisabaError classes (on @oshun/errors)
@nisaba/core depends on @oshun/errors; it has no other external or
intra-domain dependencies.
@nisaba/schemas#
Zod validation schemas for the domain entities that cross trust boundaries — for example, data arriving from external corpus connectors, user-submitted manuscripts, and API request bodies. Validated data is always preferable to raw input because Zod schemas generate both runtime validation and TypeScript types from a single source of truth.
@nisaba/schemas depends on zod and @nisaba/core. Its modules are:
manuscript.ts, variant.ts, apparatus.ts, annotation.ts, reference.ts,
tradition.ts, corpus.ts, geotemporal.ts. Each exports schemas plus
inferred types and Create* variants. See specifications.md §3 for the full
field-level breakdown.
@nisaba/languages#
The language engine handles the 29 ancient writing systems that Nisaba supports.
Each script gets its own subdirectory implementing a common ScriptHandler
interface, which ensures that higher-level code (morphological analysis,
transliteration, tokenization) can work with any writing system through the same
API surface.
The internal layout:
languages/src/
├── interfaces/
│ ├── script-handler.ts — ScriptHandler base interface
│ └── tokenizer.ts — Tokenizer interface
├── hebrew/ ancient-greek/ arabic/ aramaic/ syriac/ coptic/
├── cuneiform/ ugaritic/ old-persian/ avestan/
├── egyptian-hieroglyphic/ egyptian-hieratic/ egyptian-demotic/
├── devanagari/ grantha/ tamil-brahmi/ kharoshthi/ prakrit/
├── pali/ tibetan/ latin/ linear-b/ classical-chinese/
├── old-church-slavonic/ runic/ ethiopic/ phoenician/ samaritan/
├── tokenization/ — base tokenizer
├── transliteration/ — transliteration engine
└── lexicon/ — lexicon integration adapters
The 29 per-script directories above are the language subdirectories actually
present under languages/src/. Each typically holds handler.ts,
classifier.ts, clusters.ts, vocalization.ts, and (where applicable)
transliteration.ts and script-specific analysis files.
Every language directory implements ScriptHandler with the following
operations, all of which accept and return properly-typed Unicode text:
classify(char)— character classification within the scriptdecompose(cluster)— cluster decomposition into constituent elementsnormalize(text)— language-specific Unicode normalizationvocalize(text, scheme)— vocalization handlingtokenize(text, granularity)— position-preserving tokenizer
Key Domain Package Internals#
@nisaba/criticism#
The textual criticism engine is responsible for collating multiple manuscript witnesses, identifying and classifying variant readings, constructing stemmata, and generating critical apparatus entries. It is one of the most algorithmically complex packages in the domain because it implements established scholarly methods (Needleman-Wunsch alignment, CBGM, Lachmannian stemmatic reconstruction) rather than generic data processing.
The full set of 24 modules under criticism/src/:
| File | Purpose |
|---|---|
collation-engine.ts |
CollateX-compatible collation: Needleman-Wunsch alignment, variant-graph (DAG) construction, variant-unit detection |
fuzzy-matching.ts |
Orthographic variant detection with language-specific normalization |
block-alignment.ts |
Large-scale structural variant detection |
transposition-detection.ts |
Detection of reordered word sequences across witnesses |
apparatus-generation.ts |
Automatic apparatus entry generation from collation data |
apparatus-typography.ts |
Print-quality apparatus formatting |
cbgm-analysis.ts |
Coherence-Based Genealogical Method implementation |
lachmannian-reconstruction.ts |
Classical stemmatic reconstruction |
stemmatic-analysis.ts |
Stemma (family tree) construction and analysis |
stemma-visualization.ts |
Stemma rendering |
neighbornet.ts |
NeighborNet split-graph algorithm |
contamination-detection.ts |
Mixed manuscript tradition detection |
reconstruction-proposals.ts |
Proposed text reconstructions |
scribal-error-taxonomy.ts |
Error classification by probable scribal cause |
variant-classification.ts |
Variant unit classification |
manual-collation.ts |
Scholar-reviewed collation interface |
witness-registry.ts |
Manuscript witness registry (CRUD, filtering, fuzzy dating) |
witness-classification.ts |
Witness classification (text family, type) |
witness-hierarchy.ts |
Witness grouping and hierarchy |
siglum-management.ts |
Witness siglum registry and shorthand notation |
leiden-conventions.ts |
Leiden convention markup for epigraphic texts |
lacuna-registry.ts |
Physical damage tracking and proposed restorations |
iiif-linking.ts |
IIIF manifest integration for manuscript images |
collation-export.ts |
Export in TEI-XML, JSON, and other formats |
@nisaba/translations#
The translations package handles everything needed to compare multiple translations of the same source text — from automated statistical alignment algorithms, through neural alignment models, to multiple display modes for human readers.
translations/src/
├── alignment/ — gale-church, ibm-model1, neural-alignment, manual-alignment, anchor-index, anchor-utils
├── divergence/ — equivalence-scoring, semantic-divergence, theological-divergence, omission-detection, translation-history
├── views/ — side-by-side, synoptic, interlinear, reverse-interlinear, scroll-sync
└── types/ — segments, views type definitions
@nisaba/comparative#
The comparative package maps concepts and passages across religious and philosophical traditions. It is careful to distinguish genetic relationships (where one text directly depends on another), typological parallels (independent development of similar ideas), and contested comparisons (where scholarly opinion is divided).
| File | Purpose |
|---|---|
concept-equivalence.ts |
Curated concept mapping registry with CRUD, taxonomy, evidence grading |
parallel-passage-detection.ts |
Semantic similarity, motif matching, structural analysis |
motif-theme-tracking.ts |
Universal motif catalog with attestation and timeline |
influence-network.ts |
Directed influence graph with evidence-graded edges |
@nisaba/workspace#
The workspace package defines the scholar's primary research environment — the surface that brings together search, reading, annotations, and project management in a configurable split-pane layout.
| File | Purpose |
|---|---|
research-environment.ts |
Project lifecycle, split-pane layout, cross-corpus search |
personal-library-export.ts |
Reading lists, bookmarks, highlights, citation manager integration |
web-reading-workspace.ts |
Web reading workspace surface |
scholar-mode/ |
Scholar mode and scholar profile (scholar-profile.ts) |
Client Architecture#
Nisaba ships two distinct client packages that serve different audiences and purposes. A new engineer should understand the difference before choosing which to use.
@nisaba/api-client — generated V1 contract client#
This client is machine-generated from the OpenAPI specification and covers the
13-resource V1 contract REST API mounted at /api/v1/nisaba/*. It is the right
choice when integrating with the standard Oshun REST gateway.
api-client/src/client.ts exposes createNisabaApiClient, a thin generated
client over the OpenAPI types in api-client/src/generated/openapi.ts. It
covers list, create, get, upsert, and tombstone operations for passages,
manuscripts, editions, translations, lexicon-entries,
morphology-entries, annotations, concept-graph-nodes,
concept-graph-edges, notebooks, study-plans, citations, and
scholar-profiles. Both this file and the OpenAPI spec are generated by
libs/openapi/scripts/generate-oshun-v1-api-clients.ts.
@nisaba/client — scholarly SDK#
This is a hand-written, zero-dependency SDK (NisabaClient, built via
NisabaClientBuilder) targeting a broader scholarly API surface than the
generated V1 client. It is the right choice for application code that needs
richer scholarly operations — collation, stemma, full annotation management,
comparative concept maps, and the AI assistant.
client/src/api-client.ts covers manuscripts and witnesses, collation and
stemma, annotations, editions, canonical references, corpus search, comparative
concept maps, the AI assistant, batch requests, and WebSocket channels. It
provides automatic retry with exponential backoff, request/response
interceptors, an in-memory response cache with TTL, offset and cursor
pagination, and structured NisabaApiError errors. client/src/react-hooks.ts
adds React bindings including CRDT collaboration primitives. See
specifications.md §5 for the full SDK method/endpoint map.
The two clients do not share an endpoint surface. The V1 contract API in §4 of the specification is the generated-and-served REST surface; the
@nisaba/clientSDK targets a richer scholarly API that is not yet unified with it. The boundary exists because the V1 contract is a stable public API governed by the Oshun contract process, while the scholarly SDK is a more agile surface that can evolve as scholarly requirements develop.
Technology Stack#
The choices below reflect Nisaba's primary constraints: strict Unicode correctness for dozens of non-Latin scripts, standards compliance with TEI and W3C WADM, and interoperability with digital humanities infrastructure.
| Layer | Technology |
|---|---|
| Language | TypeScript (strict mode) |
| Runtime | Node.js |
| Testing | Vitest |
| Build | Nx with tsup/esbuild |
| Validation | Zod (@nisaba/schemas and libs/contracts/src/nisaba) |
| Database | PostgreSQL via Prisma (NISABA_DATABASE_URL) |
| API contract | OpenAPI 3.1, generated from V1 Zod contracts |
| HTTP client (SDK) | native fetch (zero dependencies) |
| Encoding standards | Unicode NFC, ISO 15924, TEI P5, W3C WADM, IIIF |
| Text protocol | CTS/CITE for canonical reference interoperability |
Adding a New Writing System#
Supporting a new ancient script requires touching four packages in order. Follow these steps to ensure that the new script integrates with all existing consumers (morphology, transliteration, search, corpus ingestion):
- Add Unicode range constants to
@nisaba/core/src/constants/unicode-ranges.ts - Register the ISO 15924 code in the
WritingSystemenum (@nisaba/core/src/types/writing-systems.ts) - Register metadata (name, direction, combining characters, ligatures) in
SCRIPT_METADATA(@nisaba/core/src/constants/script-metadata.ts) - Create a language handler directory in
@nisaba/languages/src/<script-name>/ - Implement the
ScriptHandlerinterface (languages/src/interfaces/script-handler.ts) - Register the handler in the languages index
- Write comprehensive tests covering classification, tokenization, normalization, and transliteration
Adding a New Corpus Connector#
Corpus connectors live in @nisaba/corpora, grouped by family:
classical-corpus-connectors.ts, religious-corpus-connectors.ts,
specialized-corpus-connectors.ts. These sit alongside the format parsers
(tei-xml-parser.ts, atf-parser.ts, scripture-format-parsers.ts) and the
ingestion-pipeline.ts. The key challenge when adding a new corpus is mapping
its native citation scheme to Nisaba's canonical URI format — without this
mapping, cross-corpus search and reference resolution will not work.
To add a new corpus:
- Add the connector to the relevant
*-corpus-connectors.tsmodule - Implement metadata ingestion, text retrieval, and reference mapping
- Map the corpus's native reference system to Nisaba canonical URIs
- Register the connector in the corpora index
- Write ingestion and retrieval tests
Build & Test#
- Unit tests: Vitest, per library (
*.spec.tsco-located with sources) - Integration tests:
schemas/src/__integration__/covers textual-criticism, corpus-ingestion, translation-comparison, and comparative-traditions workflows plus cross-domain (sophia-dependency) and performance benchmarks - Build: Nx with tsup/esbuild; type-check via
tsc --noEmit - CI/CD reference: A reference pipeline is documented in
docs/domains/nisaba/nisaba-ci-cd-pipeline.yaml
Scope#
This architecture document reflects libs/nisaba/* (22 packages) plus the
canonical V1 contracts in libs/contracts/src/nisaba/. Nisaba ships as
libraries and generated API contracts; there is no apps/nisaba or
services/nisaba. The scholarly-text responsibilities described here remain in
Nisaba and are not duplicated in Calliope.