Hybrid RAG workbench
An existing document chatbot rebuilt into a local evaluation workbench for classic vector retrieval and a hybrid route with query clarification, reranking and evidence compression.
A convincing chatbot answer does not establish that its retrieval found useful evidence. The original demo needed a way to inspect and compare that stage.
I revisited an older document-grounded chatbot to make retrieval inspectable. The workbench compares the classic vector-search baseline with a practical CLaRa-inspired hybrid route over the same uploaded corpus.
From requirements to working code.
- Preserved the existing upload, extraction, chunking, embedding, retrieval and cited-answer pipeline.
- Added query clarification, a broader vector candidate pool, reranking using vector and lexical signals, and evidence snippet compression.
- Built a comparison interface showing sources, retrieved chunks, candidate diagnostics and clarification behaviour.
- Added an evaluation API and CLI using expected source titles and evidence terms, with saved JSON results and a readable summary.
- Documented a phased ingestion plan that keeps local text extraction working before adding richer table and figure handling.
Keep a baseline to compare against
Classic retrieval remains available alongside the hybrid route so changes can be judged against the same documents and questions.
Inspect evidence separately from answers
Source matches, term coverage, retrieval latency and compression expose different trade-offs. A fluent generated response is not treated as an evaluation result.
Build locally before adding providers
Offline mock mode supports local development. Structured extraction, newer Azure document processing and selective visual analysis are documented next steps.
What the work supports.
- Classic and hybrid chat routes, the comparison interface and the labelled evaluation API and CLI are present in the local workspace.
- The hybrid service exposes its vector, lexical and query-term coverage signals along with compression diagnostics.
- Evaluation outputs retain the per-case results rather than only report aggregate scores.
The current workbench is local development and is not represented by the older hosted demo. This is inspired by CLaRa, not a reproduction of its full architecture. No measured retrieval improvement is claimed. Structured table and figure ingestion and the newer Azure provider remain planned.