Skip to content
AI & RAG

Why most company chatbots hallucinate — and how we built one that cites its sources

A behind-the-scenes look at AIVCJ Knowledge: hybrid retrieval, query planning, a grounding check on every answer, and two public test sets — including a blind one written the way real people type. Try it live on sample HR, product and GST documents, or your own PDF.

4 min read
Answers that cite the exact passage

Every company has the same problem. The answers exist — in a leave policy, a warranty PDF, an SOP from 2023 — but nobody can find them, so people ask a colleague, guess, or open a ticket. A chatbot looks like the obvious fix. Most of them disappoint for one reason: they answer confidently even when they don't know, and nobody can check where an answer came from.

We built AIVCJ Knowledge to show what the alternative looks like. It is a live demo you can open today: pick a sample library, ask a real question, and every sentence in the answer points to the exact passage it came from.

The problem: fluent is not the same as correct

A language model on its own answers from memory. Ask it about your carry-forward rule and it will produce something plausible from thousands of other companies' policies. Plain retrieval (RAG) helps, but naive versions fail in predictable ways:

  • Multi-part questions. "I joined two months ago in Finance — can I work from home two days a week?" needs the hybrid policy and the probation rule. One search query rarely finds both.
  • Tables and headings get split. Cut a document every 900 characters and the allowance table ends up separated from the heading that says what it is.
  • No honest "I don't know". When the right passage isn't retrieved, the model fills the gap.
  • No way to check. Without citations, a wrong answer looks exactly like a right one.

How AIVCJ Knowledge works

  1. Structure-aware ingestion. Documents (PDF, DOCX, TXT) are split along their own structure: headings stay with their content, tables stay whole, and every passage remembers its document and section. That context is embedded with the passage, so a table row still "knows" it belongs to the travel policy's allowance section.
  2. Query planning. A small, fast model splits each question into the separate facts it needs — the employee's grade, the city's tier, the allowance table — and each becomes its own search.
  3. Hybrid retrieval. Every search runs two ways inside PostgreSQL: meaning (pgvector embeddings) and exact words (full-text search). Results are merged with reciprocal-rank fusion, and each part of the question is guaranteed a slot.
  4. Answer only from the sources. Claude writes the answer from the retrieved passages alone, citing each claim as [1], [2]… If the documents don't cover it, it says so plainly.
  5. Grounding check. A second pass compares the finished answer with its sources and labels it fully supported, partially supported or not in the documents — visible to the user.

What you see in the demo

  • Clickable citations that open the source page with the exact passage highlighted.
  • "Why this answer?" — every retrieved passage with its meaning, keyword and fusion scores, which ones were cited, the grounding verdict, time taken and tokens used.
  • Your own document. Upload a PDF or Word file (private to you, deleted nightly) and ask questions about it within seconds.
  • A quality page with the full test results, question by question.

Measured, not guessed — and honest about the limits

Each sample library has two test sets, and every answer is checked for faithfulness to the sources and relevance to the question.

  • Golden set (45 questions). Written alongside the documents, including multi-document questions and questions the documents deliberately don't answer. We used it while tuning, so it flatters us. Our first version — fixed-size chunks and a single search per question — scored 40 of 45 with 91–93% faithfulness; the misses were all multi-part questions. Structure-aware chunking, contextual embeddings and query planning lifted it to 42–45 of 45 across runs, with 98–100% faithfulness.
  • Blind set (20 questions). Written afterwards by a different author the way people really type — Hinglish, typos, vague wording, questions that need two documents — and never used for tuning: 19–20 of 20 across runs, with 95–96% faithfulness. The miss in one run is instructive: asked (in Hinglish) what to do after leaving a company laptop in a cab, it gave the regular IT helpdesk instead of the 24×7 IT Security line and skipped the police-report step — exactly the kind of gap a blind set exists to catch.

Two caveats we'd rather state than hide: the grading is done by an AI judge (Claude) against reference answers, the same model family as the assistant, and the results are not human-verified; and our sample documents are cleaner than most real company files. Treat the numbers as indicative. Every question, answer and grade is public on the quality page.

Just as important: questions the documents don't cover get an honest "I couldn't find this in the documents" — not an invention.

This is the discipline we bring to client projects: a test set of real questions from day one, re-run whenever a prompt, model or document changes, so quality is a number you can track rather than a feeling.

Running it on your data

The demo runs on the same stack we deploy for clients: Next.js, PostgreSQL with pgvector, and Claude, on infrastructure you choose — including your own servers for sensitive documents. A typical engagement starts with one high-volume workflow (HR questions, product support or internal SOPs), fifty real questions, and a measured pilot in weeks, not months. Add connectors (Google Drive, SharePoint, Notion), per-document permissions and a WhatsApp channel as you grow.

Try AIVCJ Knowledge live →

  • #RAG
  • #Case study
  • #pgvector
  • #Claude
  • #Evaluation
LinkedInWhatsApp
Frequently asked questions

Frequently asked questions

01Can it work on our own documents?
Yes. The demo lets you upload your own PDF or Word file privately. For a production system we connect your drives and knowledge bases and respect who can see which document.
02What happens when the answer isn't in the documents?
It says so. Answers are written only from retrieved passages, and a grounding check labels every answer as supported, partially supported or not found.
03Can it be hosted on our own servers?
Yes. The stack is Next.js and PostgreSQL with pgvector, and it can run on your servers or private cloud for sensitive data.
Need this built?

AI & RAG Applications

Assistants that answer from your data — and cite it.

Keep reading