Why most company chatbots hallucinate — and how we built one that cites its sources
A behind-the-scenes look at AIVCJ Knowledge: hybrid retrieval, query planning, a grounding check on every answer, and two public test sets — including a blind one written the way real people type. Try it live on sample HR, product and GST documents, or your own PDF.

Every company has the same problem. The answers exist — in a leave policy, a warranty PDF, an SOP from 2023 — but nobody can find them, so people ask a colleague, guess, or open a ticket. A chatbot looks like the obvious fix. Most of them disappoint for one reason: they answer confidently even when they don't know, and nobody can check where an answer came from.
We built AIVCJ Knowledge to show what the alternative looks like. It is a live demo you can open today: pick a sample library, ask a real question, and every sentence in the answer points to the exact passage it came from.
The problem: fluent is not the same as correct
A language model on its own answers from memory. Ask it about your carry-forward rule and it will produce something plausible from thousands of other companies' policies. Plain retrieval (RAG) helps, but naive versions fail in predictable ways:
- Multi-part questions. "I joined two months ago in Finance — can I work from home two days a week?" needs the hybrid policy and the probation rule. One search query rarely finds both.
- Tables and headings get split. Cut a document every 900 characters and the allowance table ends up separated from the heading that says what it is.
- No honest "I don't know". When the right passage isn't retrieved, the model fills the gap.
- No way to check. Without citations, a wrong answer looks exactly like a right one.
How AIVCJ Knowledge works
- Structure-aware ingestion. Documents (PDF, DOCX, TXT) are split along their own structure: headings stay with their content, tables stay whole, and every passage remembers its document and section. That context is embedded with the passage, so a table row still "knows" it belongs to the travel policy's allowance section.
- Query planning. A small, fast model splits each question into the separate facts it needs — the employee's grade, the city's tier, the allowance table — and each becomes its own search.
- Hybrid retrieval. Every search runs two ways inside PostgreSQL: meaning (pgvector embeddings) and exact words (full-text search). Results are merged with reciprocal-rank fusion, and each part of the question is guaranteed a slot.
- Answer only from the sources. Claude writes the answer from the retrieved passages alone, citing each claim as [1], [2]… If the documents don't cover it, it says so plainly.
- Grounding check. A second pass compares the finished answer with its sources and labels it fully supported, partially supported or not in the documents — visible to the user.
What you see in the demo
- Clickable citations that open the source page with the exact passage highlighted.
- "Why this answer?" — every retrieved passage with its meaning, keyword and fusion scores, which ones were cited, the grounding verdict, time taken and tokens used.
- Your own document. Upload a PDF or Word file (private to you, deleted nightly) and ask questions about it within seconds.
- A quality page with the full test results, question by question.
Measured, not guessed — and honest about the limits
Each sample library has two test sets, and every answer is checked for faithfulness to the sources and relevance to the question.
- Golden set (45 questions). Written alongside the documents, including multi-document questions and questions the documents deliberately don't answer. We used it while tuning, so it flatters us. Our first version — fixed-size chunks and a single search per question — scored 40 of 45 with 91–93% faithfulness; the misses were all multi-part questions. Structure-aware chunking, contextual embeddings and query planning lifted it to 42–45 of 45 across runs, with 98–100% faithfulness.
- Blind set (20 questions). Written afterwards by a different author the way people really type — Hinglish, typos, vague wording, questions that need two documents — and never used for tuning: 19–20 of 20 across runs, with 95–96% faithfulness. The miss in one run is instructive: asked (in Hinglish) what to do after leaving a company laptop in a cab, it gave the regular IT helpdesk instead of the 24×7 IT Security line and skipped the police-report step — exactly the kind of gap a blind set exists to catch.
Two caveats we'd rather state than hide: the grading is done by an AI judge (Claude) against reference answers, the same model family as the assistant, and the results are not human-verified; and our sample documents are cleaner than most real company files. Treat the numbers as indicative. Every question, answer and grade is public on the quality page.
Just as important: questions the documents don't cover get an honest "I couldn't find this in the documents" — not an invention.
This is the discipline we bring to client projects: a test set of real questions from day one, re-run whenever a prompt, model or document changes, so quality is a number you can track rather than a feeling.
Running it on your data
The demo runs on the same stack we deploy for clients: Next.js, PostgreSQL with pgvector, and Claude, on infrastructure you choose — including your own servers for sensitive documents. A typical engagement starts with one high-volume workflow (HR questions, product support or internal SOPs), fifty real questions, and a measured pilot in weeks, not months. Add connectors (Google Drive, SharePoint, Notion), per-document permissions and a WhatsApp channel as you grow.
- #RAG
- #Case study
- #pgvector
- #Claude
- #Evaluation
Frequently asked questions
01Can it work on our own documents?
02What happens when the answer isn't in the documents?
03Can it be hosted on our own servers?
AI & RAG Applications
Assistants that answer from your data — and cite it.
Keep reading
ERP & Operations5 min read
From chatbot to co-worker: building an HR assistant that takes actions safely
AIVCJ HR doesn't just answer leave questions — it applies leave, routes approvals and drafts HR letters. Here is how we let an AI act without letting it make the rules: a policy engine in code, a confirmation card before every action, role permissions and an audit log. Measured on a scripted and a blind test set.
Read more
ERP & Operations6 min read
From WhatsApp photos to a live order book: an ERP and offline field app for distributors
AIVCJ ERP replaces the Tally + Excel + WhatsApp loop: a field app that takes orders without signal and never syncs twice, credit holds decided in code, FEFO picking with GST invoices, and plain-English questions answered by a read-only query you can see. Measured with a two-phone sync test and a blind NL-to-SQL set.
Read more2 min read
What is RAG — and when does your business actually need it?
Retrieval-augmented generation lets AI answer from your own documents instead of guessing. Here is how it works, where it shines, and the signs you are ready for it.
Read more