Tessmora · Sources to answers
Modality-aware retrieval, a separate Pi Agent, and verifiable citations.
What it does
Tessmora is a self-hostable knowledge platform for asking questions across documents, images, audio, and video. It keeps modality-appropriate retrieval units and returns answers that can be traced to source passages or media locations.
The system combines a React interface, a Python/FastAPI host, Qdrant indices, and MinIO originals. Ordinary direct retrieval, a legacy bounded Agent loop, and the separate Pi Agent path have different execution responsibilities.
Section sources
Modality-aware ingestion
Documents retain source-faithful atomic units; images combine descriptions and CLIP features; audio combines speech transcription and CLAP; video centers retrieval on semantic shots with captions, speech, and supporting keyframes. Originals and retrieval representations serve different purposes.
The document range planner proposes ranges over original units. Validation requires contiguous, non-overlapping, complete coverage and a size bound; invalid plans fall back to deterministic splitting. Model-generated boundaries cannot silently drop or rewrite source text.
Retrieval and scope
Retrieval combines dense, sparse, and media-specific recall, then fuses and reranks candidates. Knowledge-base portraits can help choose a scope when it is not pinned. Permissions and the visitor's selected search scope constrain which sources can participate.
An @ reference is input material, not automatically the search scope. Pi's immutable AccessScope distinguishes allowed knowledge bases, search knowledge bases/files, and input references. A bound input reference should not be presented as a newly retrieved search result.
The Pi Agent path
Pi Agent is a separate integration of the earendil-works/pi loop in a Node runtime with a Python host. It is not a new label for the older bounded Agent mode. Pi owns iterative research, source selection, and final writing; the host owns scope, permissions, source identity, usage, events, and final validation.
Its tools expose scoped, read-only retrieval and evidence operations. They do not grant a general shell, arbitrary filesystem access, or unrestricted web tools. Task completion, cancellation, and errors control the current Pi lifecycle; the old loop's fixed round limits must not be generalized to every mode.
Evidence and citations
Pi stores runs, events, and evidence in SQLite. Evidence identity hashes the source, version, locator, content, and observation; repeated evidence is deduplicated and gets a stable evidence number. Terminal runs cannot simply be rewritten.
Citations connect an answer to document passages, images, audio, or video locations. Inspecting a locator lets a reader check what supports a claim. A displayed citation or an architecture illustration alone does not prove that an answer is correct.
Access and limits
The repository is public; no software license file was found at review time. Self-hosting needs configured parsing and model services plus data stores. The default local/trusted-network setup is not a ready-made public multi-tenant service; public exposure needs authentication and operational controls.
Current retrieval evaluation work in the local checkout is still being edited. This Wiki makes no new accuracy, latency, cost, or competitor-superiority claims based on unfinished experiments or test-count descriptions.