This is Anil's live RAG.
Built to show the retrieval layer, not a landing-page chatbot. I can audit your documents or install the same class of system for your team.
anilpervaiz.comHybrid retrieve, cite the page, verify the claim, refuse when the file is silent. A production RAG stack you can hire, not a chat wrapper.
CiteRAG is the working case: ingest, hybrid retrieve, cite, verify, refuse, evaluate. The offer is a demo, an accuracy audit, or a custom pipeline. Not a $5 SaaS seat.
PDF, DOCX, MD, TXT. Semantic split. Source, page, and section on every chunk. That is what a citation is made from.
Vector finds meaning. BM25 finds the clause number. Reciprocal rank fuses them. Cohere rerank is optional.
The model must cite. A second pass asks if the chunk supports the sentence. If the file is silent, the system says so.
Vector-only RAG falls off the keyword side. Keyword-only misses meaning. Production RAG runs both, then lets the answer flow from the fuse.
A citation is not decoration. The river is the claim: from sentence, to [Source: file, Page X], to a judge that checks the chunk. If the water stops, we refuse.
Prototype RAG is embed-and-chat. Production RAG is hybrid retrieve, rerank, cite, measure, and trace.
| Industry baseline | In CiteRAG | |
|---|---|---|
| Semantic chunking with source metadata | SemanticSplitter. source, page, section, indexed_at | src/ingestion |
| Hybrid search: dense + BM25, fused with RRF | QueryFusionRetriever mode reciprocal_rerank | src/retrieval |
| Cross-encoder rerank | CohereRerank when a key is set | optional |
| Grounded citations, not prompt hope | Forced cite + claim judge + refuse | src/generation |
| RAGAS + a CI-style gate | faithfulness, relevancy, precision, recall. DeepEval. audit-report. | src/evaluation |
| Query traces | Langfuse + structured logs | src/utils |
| HTTP API, auth split, persistence option | FastAPI. Demo key vs API key. Memory or Pinecone. | src/api |
Baseline from 2026 RAG architecture writeups (hybrid + RRF + rerank, RAGAS, traces). No live faithfulness KPI on this page. Render can be cold. Cohere and Pinecone are env keys.
Nothing here is a hypothetical roadmap. Column three is what the repo does today; column four is the honest gap. The demo you clicked runs tier one, on the live API.
| Scale | What has to change | In the repo today | Gap to close |
|---|---|---|---|
| Tier 1 Small team ~1k docs |
One process, in-memory index, shared demo key | Live now: memory backend, BM25 + vector fusion, demo key | — |
| Tier 2 Department ~50k docs |
Index outlives the process; per-team keys | Pinecone store coded behind use_pinecone + key | Needs a real Pinecone index; keys are still a single shared secret |
| Tier 3 Company multi-team |
Rate limits and auth survive more than one replica | API key vs demo key split, sliding-window rate limit, audit trail via traces | Rate limiter is in-memory — per instance, not shared across replicas |
| Tier 4 Regulated org |
Identity, tenancy, residency, sign-off | Structured logs, per-query decision records with cost and confidence | No SSO/SAML, no multi-tenancy, no RBAC built. Sold as custom work, not shipped features |
Verifying a claim used to mean asking a frontier model to "reply yes or no" — once per claim. On a reasoning model that is roughly a thousand hidden tokens per boolean.
Now every claim in an answer is checked in one batched decision call that returns a calibrated probability instead of a bare yes. Measured live on this demo: 61 claims across 17 queries for $0.0019 total, at 100% verification accuracy.
Report: docs/reports/2026-09-26_131644_live. Real run, not a projected figure.
Before generation, a typed gate scores the query. On the 17-question golden set the separation was total — no overlap at all:
It also routed the one genuinely multi-document question to hybrid search and scored it "needs synthesis across documents" — without being told which questions were multi-hop. Exact-number queries went to keyword search; prose questions went to vector.
A fabricated claim contradicted by its own source scored 0.01. Decisions are calibrated probabilities, never a guarantee — nothing irreversible is auto-executed.
An accuracy audit is the same path as the stack: connect documents, measure the baseline, fix retrieval, measure again. You leave with a score you can rerun, not a slide.
Built to show the retrieval layer, not a landing-page chatbot. I can audit your documents or install the same class of system for your team.
anilpervaiz.comThree-pane UI on the sample corpus. Ask, see the cite, see verify.
Open /demoBaseline on your files, then a fix list. Mail hello@anilpervaiz.com.
Book an auditLangChain is a kit. This is a finished citation path on LlamaIndex: hybrid retrieve, forced cites, claim verification, refusal, eval.
No. Render can cold-start. Local uvicorn is the reliable path.
No. Documents stay in the index you run.
Yes. Most work is the retrieval layer and verify, not a rewrite.
CICADA original stays at /v1.