What the system returns
A user asks a plain-language question and receives an answer with a document title, page, sheet or slide. Complex questions are split into multiple retrieval topics. Missing evidence for one part produces a clarification or refusal. A client report uploaded to the system does not automatically become a normative source.
Document pipeline
- 01Download the file and preserve its source URL, MIME type, size and SHA-256.
- 02Extract PDF, Word, Excel or PowerPoint structure; run OCR only for weak pages.
- 03Create chunks with page, sheet, cell range or slide addresses.
- 04Build lexical and semantic indexes.
- 05Generate only from retrieved evidence and validate citations on the server.
Why retrieval uses two channels
FTS5/BM25 handles order numbers, references and exact terms. Embeddings retrieve paraphrases. Reciprocal Rank Fusion and metadata filters combine both rankings.
A complex question becomes two to four independent subqueries. Missing evidence for one step blocks a complete answer.
exact terms → FTS5 / BM25
plain-language question → embeddings
rankings → RRF → evidence set
claims → citation validation → answerPreventing invented citations
The model may select only known chunk IDs. The server verifies that each ID exists, belongs to the stated document and contains the quoted text. An unknown citation, an uncited claim or a changed number forces a refusal.
Document text is treated as untrusted data. An instruction embedded in a PDF cannot change system rules or request a secret.
Numbers and document currency
A narrow water-withdrawal calculation runs in deterministic Decimal code. The model identifies explicit inputs, while the formula, substitution and result remain auditable.
A document being present in an official catalog does not prove that the edition is currently valid. That requires a separate status registry and verification source. A domain specialist still confirms critical conclusions.