Skip to main content
All projects
Agentic RAG

Noesis

A private research assistant that answers from your own English and Arabic documents, checks every answer against its sources, and turns what it reads into training data.

Try it livePrivate sourceWatch the run
Noesis
  • 93.9%

    correct passage ranked first

  • 54 ms

    median search over ~97,000 passages

  • 90%

    right document found across languages

  • 3.3 s

    Arabic answer from English sources, fully verified

Inside the product

See it working

Real screens and a recorded run, captured from the working product.

Noesis
A real session: sign in, the overview, a new question typed and answered live with citations and verification, the research steps and a citation preview, then the knowledge graph with a document opened.
Noesis
Open full size

Sign-in page with the animated particle knowledge sphere and the headline "Turn your documents into an expert that cites its sources."

01 / 13

My role

Sole engineer: retrieval, the tool-using assistant, the citation checker, the data pipeline, the security model and the bilingual interface.

The problem

Teams sit on large private collections in two languages, and a chatbot that is fed whole documents is slow, expensive and happy to make things up. People need answers they can trust, with the exact passage behind every claim, from data that never leaves their machine, and without running a database server or Docker.

What I built

A full-stack knowledge platform that runs from one folder. Documents (PDF, Word, Excel, CSV, HTML, Markdown, JSON) are split into passages and indexed three ways: SQLite FTS5 with BM25, Gemini embeddings in an on-disk HNSW index, and local hashed vectors that work offline. The signals are fused with reciprocal-rank fusion and diversified with MMR, so an Arabic question finds an English source and the other way round.

The assistant never receives the documents. It gets the question and six tools (search, read a passage, outline a document, find related material, check reviewed Q&A, survey the collection), decides what to look up and stops when the evidence is enough. Every answer is then checked sentence by sentence against the passages it cites and labelled verified, cited across languages, partially verified or not verified.

Architecture

Python FastAPI backend on SQLite (WAL and FTS5) with USearch HNSW indexes on disk, all in one data folder, so a backup is a folder copy. A React 19, Vite and Tailwind 4 frontend with TanStack Query and Virtual, Zustand, Motion and a d3-force knowledge graph laid out in a web worker. Around the assistant sit a knowledge graph that links documents by meaning, a miner that drafts grounded question and answer pairs and keeps only those that pass quality checks, a dataset manager that exports JSONL, chat messages, Alpaca, CSV and Excel, and an admin area with accounts, page-level permissions, audit trail and model usage with cost.

Challenges

Answering across languages without translating the collection, keeping search fast on about 100,000 passages on an ordinary machine, and proving each answer without paying for a second model call. The citation check compares wording, word forms, numbers and names, even between Arabic and English, and runs instantly and deterministically.

Key engineering decisions

Tools instead of stuffing context: the model sees only what it asks for, about 3,000 to 3,500 input tokens per answer. Weak matches are labelled weak, so a question the collection cannot answer is declined instead of guessed. Caching of repeated questions tied to the data, model and instructions. A full offline mode that still searches and answers with quoted passages when no key is set.

Results & impact

Measured with the project's own benchmark: the correct passage is ranked first 93.9% of the time with embeddings (85.6% offline) and is in the top five 100% of the time; a document in the other language is ranked first 90% of the time. Search over about 97,000 passages takes a median 54 ms (95th percentile 78 ms) and a dense vector query about 2 ms. In the recorded runs, a two-part English question was answered in 9.6 s after 2 research steps and 6 sources, verified 100%; an Arabic question was answered in Arabic from English sources in 3.3 s; two answers cost about $0.001 in total. The miner turned 2 passages into 5 grounded Q&A pairs, each scored 100%.

Highlights

  • Tool-using research assistant with streamed answers, numbered citations and a hover preview of each cited passage
  • Deterministic citation check on every answer, including across English and Arabic
  • Hybrid bilingual retrieval: BM25, Gemini embeddings (HNSW) and local vectors fused with RRF and MMR
  • Retrieval lab that shows which signal ranked each passage
  • Knowledge graph linking documents by meaning across languages
  • Miner and dataset manager for grounded training data, exported as JSONL, chat messages, Alpaca, CSV or Excel
  • Argon2id passwords, rotating refresh tokens with reuse detection, lockout, rate limits and CSP
  • No Docker, no database server: start.cmd or start.sh installs and runs everything
  • English and Arabic UI with full RTL, light and dark themes, Ctrl+K and an interactive particle knowledge sphere