Your AI answers.
Proven by citations.

Hybrid retrieve, cite the page, verify the claim, refuse when the file is silent. A production RAG stack you can hire, not a chat wrapper.

Documents
refund-policy.pdf
terms-2026.pdf
compliance.md
pricing.docx
Chat
What is the refund policy for annual subscriptions?
Annual subscriptions can be refunded within 30 days of purchase. Source: refund-policy.pdf, p.2
Verification
refund-policy.pdfsupported
not in corpuswould refuse
Sample faithfulness
0.94

I built this stack.
I can build it on your documents.

CiteRAG is the working case: ingest, hybrid retrieve, cite, verify, refuse, evaluate. The offer is a demo, an accuracy audit, or a custom pipeline. Not a $5 SaaS seat.

01

Ingest for real.

PDF, DOCX, MD, TXT. Semantic split. Source, page, and section on every chunk. That is what a citation is made from.

02

Retrieve twice.

Vector finds meaning. BM25 finds the clause number. Reciprocal rank fuses them. Cohere rerank is optional.

03

Then prove it.

The model must cite. A second pass asks if the chunk supports the sentence. If the file is silent, the system says so.

Two cliffs. One stream.

Vector-only RAG falls off the keyword side. Keyword-only misses meaning. Production RAG runs both, then lets the answer flow from the fuse.

  • VectorIndexRetriever semantic neighbors
  • BM25Retriever exact tokens, SKUs, clause IDs
  • QueryFusionRetriever reciprocal_rerank

Follow the claim downstream.

A citation is not decoration. The river is the claim: from sentence, to [Source: file, Page X], to a judge that checks the chunk. If the water stops, we refuse.

  • Forced cite prompt requires the marker
  • Claim judge LLM check, substring fallback
  • Honest refuse when the documents do not say

What production RAG requires.
What this repo ships.

Prototype RAG is embed-and-chat. Production RAG is hybrid retrieve, rerank, cite, measure, and trace.

Industry baselineIn CiteRAG
Semantic chunking with source metadataSemanticSplitter. source, page, section, indexed_atsrc/ingestion
Hybrid search: dense + BM25, fused with RRFQueryFusionRetriever mode reciprocal_reranksrc/retrieval
Cross-encoder rerankCohereRerank when a key is setoptional
Grounded citations, not prompt hopeForced cite + claim judge + refusesrc/generation
RAGAS + a CI-style gatefaithfulness, relevancy, precision, recall. DeepEval. audit-report.src/evaluation
Query tracesLangfuse + structured logssrc/utils
HTTP API, auth split, persistence optionFastAPI. Demo key vs API key. Memory or Pinecone.src/api

Baseline from 2026 RAG architecture writeups (hybrid + RRF + rerank, RAGAS, traces). No live faithfulness KPI on this page. Render can be cold. Cohere and Pinecone are env keys.

The same pipeline,
at four scales.

Nothing here is a hypothetical roadmap. Column three is what the repo does today; column four is the honest gap. The demo you clicked runs tier one, on the live API.

ScaleWhat has to changeIn the repo todayGap to close
Tier 1
Small team
~1k docs
One process, in-memory index, shared demo key Live now: memory backend, BM25 + vector fusion, demo key —
Tier 2
Department
~50k docs
Index outlives the process; per-team keys Pinecone store coded behind use_pinecone + key Needs a real Pinecone index; keys are still a single shared secret
Tier 3
Company
multi-team
Rate limits and auth survive more than one replica API key vs demo key split, sliding-window rate limit, audit trail via traces Rate limiter is in-memory — per instance, not shared across replicas
Tier 4
Regulated org
Identity, tenancy, residency, sign-off Structured logs, per-query decision records with cost and confidence No SSO/SAML, no multi-tenancy, no RBAC built. Sold as custom work, not shipped features

Decisions left the expensive model.

Verifying a claim used to mean asking a frontier model to "reply yes or no" — once per claim. On a reasoning model that is roughly a thousand hidden tokens per boolean.

Now every claim in an answer is checked in one batched decision call that returns a calibrated probability instead of a bare yes. Measured live on this demo: 61 claims across 17 queries for $0.0019 total, at 100% verification accuracy.

Report: docs/reports/2026-09-26_131644_live. Real run, not a projected figure.

It knows when it cannot answer.

Before generation, a typed gate scores the query. On the 17-question golden set the separation was total — no overlap at all:

  • Unanswerable questions: 0.01 – 0.03
  • Answerable questions: 0.88 – 0.98

It also routed the one genuinely multi-document question to hybrid search and scored it "needs synthesis across documents" — without being told which questions were multi-hop. Exact-number queries went to keyword search; prose questions went to vector.

A fabricated claim contradicted by its own source scored 0.01. Decisions are calibrated probabilities, never a guarantee — nothing irreversible is auto-executed.

Walk your corpus uphill.

An accuracy audit is the same path as the stack: connect documents, measure the baseline, fix retrieval, measure again. You leave with a score you can rerun, not a slide.

  • 01 Connect your files, or start on the demo corpus
  • 02 Measure RAGAS on real questions
  • 03 Leave with a before / after you can defend

Three ways in.
No fake pricing grid.

This is Anil's live RAG.

Built to show the retrieval layer, not a landing-page chatbot. I can audit your documents or install the same class of system for your team.

anilpervaiz.com

Start demo

Three-pane UI on the sample corpus. Ask, see the cite, see verify.

Open /demo

Book an audit

Baseline on your files, then a fix list. Mail hello@anilpervaiz.com.

Book an audit

Questions.

How is this different from LangChain?

LangChain is a kit. This is a finished citation path on LlamaIndex: hybrid retrieve, forced cites, claim verification, refusal, eval.

Is the hosted API always hot?

No. Render can cold-start. Local uvicorn is the reliable path.

Does my data train a model?

No. Documents stay in the index you run.

Can you use our existing stack?

Yes. Most work is the retrieval layer and verify, not a rewrite.

Where is the original landing?

CICADA original stays at /v1.

Get answers you can prove.

Start demo Book an audit