Omni-modal retrieval, one traceable evidence chain
How can each modality keep an appropriate retrieval path and still join one inspectable answer?
Project context
Tessmora sends documents, images, audio, and video through dedicated parsing and embedding paths. One-Pass intent and knowledge-base portraits choose the search scope; dense, sparse, visual, audio, and video retrieval then feed weighted RRF, Cross-Encoder reranking, and either a direct answer or a budget-bounded read-only evidence loop.
SYSTEM PATH
Distinct paths, one cited answer
01Modality-aware ingestion
02KB portrait routing
03Five-route retrieval and reranking
04Direct or bounded evidence loop
05Typed citations over SSE
Key decisions
Keep a dedicated path for each modality
Documents use source-faithful agentic chunks; images use VLM descriptions and CLIP; audio uses ASR and CLAP; video is indexed around semantic shots with aligned captions, speech, and key-frame visual support. The goal is to avoid collapsing every source into one text-only representation.
Route scope before ranking
When a knowledge base is not pinned, topic portraits can route a query to one, several, or all relevant bases. Dense, sparse, visual, audio, and video candidates are fused by weighted RRF and narrowed by Cross-Encoder reranking.
Bound the evidence loop
Agent mode plans complementary subqueries and executes each round in parallel through the same read-only retriever. Round, query, and evidence budgets bound the loop; the final stream sends type-specific citations separately from answer text.
DESIGN CONTRACT
Principles and honest boundaries
Do not flatten modalities
Bind answers to evidence
Give agents explicit budgets
Keep self-hosting boundaries clear
Implementation and visuals
Project UI, architecture, and system visualizations used to make the work inspectable.
Knowledge-base portraits, five retrieval routes, fusion, reranking, and direct or bounded answer paths
Type-specific document, image, audio, and video locators attached to one answer