Skip to content
ERP & Operations

From chatbot to co-worker: building an HR assistant that takes actions safely

AIVCJ HR doesn't just answer leave questions — it applies leave, routes approvals and drafts HR letters. Here is how we let an AI act without letting it make the rules: a policy engine in code, a confirmation card before every action, role permissions and an audit log. Measured on a scripted and a blind test set.

5 min read
Applies leave within the rules — you confirm

Ask any HR team where their week goes and you hear the same list: “How many leaves do I have?”, leave forms on WhatsApp, chasing managers for approvals, and “Can you send me an experience letter?”. A policy chatbot helps with the first one. It does nothing for the rest — and the rest is where the hours go.

We built AIVCJ HR to show the next step: an assistant that answers and acts. It checks your balance, applies leave within the rules, routes it to your manager with a summary, and drafts HR letters — inside a sample company, Northwind Distribution, with 42 employees across a head office and three warehouses.

Why answers alone don't reduce HR load

Our Knowledge demo answers policy questions with citations. That removes the “what does the policy say?” question. But the employee still has to open a form, count working days, check the notice period, find the approver and wait. The work moved; it didn't disappear.

Letting an AI do the work raises a fair worry: what stops it from approving its own mistakes, bending a rule to be helpful, or showing one employee another's data? So we designed the guard-rails first.

The guard-rails

  1. The rules live in code, not in the model. A deterministic policy engine decides how many days a request counts (weekly offs and holidays per location, the sandwich rule for earned leave), whether the notice period is met, whether the employee is still on probation, whether the balance covers it, and who must approve. The AI calls the engine and explains the result — it cannot overrule it.
  2. Nothing happens without a click. Every write — applying leave, cancelling a request, raising an IT or asset request — is shown as a confirmation card with the counted days, every check that passed or failed, the balance before and after, and the approval route. The request is created only when the person presses Confirm, and the engine re-checks it at that moment.
  3. Role permissions are enforced by the tools. An employee sees only their own data; a manager sees their direct reports; only HR can draft letters for others. Approving is never done by the AI — managers decide on the Approvals screen.
  4. Everything is audited. Every tool call — who, which role, what the AI asked for, and whether it ran, was refused by the rules, or was confirmed by a human — lands in an audit log on the HR dashboard.

What you can try in the demo

  • As Priya (employee): “I need leave from 20 to 24 Oct for my sister's wedding.” The engine notices that 20 Oct is a holiday and 24 Oct a weekly off, counts 3 days, explains why marriage leave doesn't apply to a sibling's wedding, and shows the card.
  • As Rohit (manager): the request is waiting with an AI summary — within policy, Sneha is out on 21–22 Oct, earned leave 12 → 9. Approve it, and Priya's balance and request timeline update live.
  • As Anjali (HR): “Experience letter for Karan Verma” — drafted from his record, edited in the browser, downloaded as a PDF. Anything the record doesn't contain is left as a placeholder rather than invented.
  • Policy questions (“Can I claim my Uber to a client meeting?”) use the same cited retrieval engine as AIVCJ Knowledge.

Every visitor gets their own copy of the company, so nobody's approvals collide with anyone else's, and it resets every night.

Measured — including where it slipped

We test the assistant end to end — same model, tools, engine and permissions as the live demo — on two sets of requests, each in a fresh copy of the company. A case passes only if the assistant ends with exactly the right action (or correctly takes none), and, where a reference answer exists, an AI judge agrees with the reply.

  • Scripted set (30 requests), written with the app and used while tuning, including 12 that break a rule or a permission: 30 of 30 correct on the final run, with all 12 violations refused (earlier runs while tuning: 29 of 30).
  • Blind set (20 requests), written afterwards by a different author who never saw our prompts or code — Hinglish, typos, vague dates (“parso se 3 din”), two asks in one message: 17 of 20 (85%) in both runs.

The most useful results were the failures. In one early run an employee asked for an experience letter and the assistant, trying to help, wrote one in the chat — exactly the kind of document only HR should issue. In another it asked “shall I file this?” instead of running a 4-day casual-leave request through the engine, so it never said the limit is 3 days. Both became explicit rules. The blind set found three more. An employee asked for five days of casual leave: the engine refused it correctly (the limit is three), but the assistant then prepared an earned-leave card on its own instead of asking first — nothing was submitted without a click, but it shouldn’t have switched the leave type. A manager asked for “neha”’s balance and the name search also matched “Sneha”, so it asked a needless question. The third was a correct refusal that the AI judge marked down for leaving out an extra detail. We fixed the first two; because those fixes came from the blind set, we don’t count re-runs after them as blind.

Two caveats we'd rather state than hide: the judging is done by an AI (the same model family as the assistant) and is not human-verified, and a test set of 50 requests is small. Treat the numbers as indicative. What the guard-rails guarantee is structural: whatever the model says, a request that breaks a rule cannot be submitted, because the engine — not the model — decides.

Deploying on your HRMS

The demo runs on the stack we deploy for clients: Next.js, PostgreSQL with pgvector, and Claude. In a real rollout the policy engine is configured from your leave policy, the tools talk to your HRMS (Keka, Darwinbox, Zoho People or your own), and the same assistant can sit on WhatsApp. A typical engagement starts with leave and letters for one location, measured on fifty real employee requests, before expanding.

Try AIVCJ HR live →

  • #AI agents
  • #HR tech
  • #Case study
  • #Claude
  • #Evaluation
LinkedInWhatsApp
Frequently asked questions

Frequently asked questions

01Can the AI approve or submit anything on its own?
No. The AI prepares a confirmation card; a person must press Confirm, and approvals are always made by the manager on the Approvals screen. Every step is in the audit log.
02What if the AI misunderstands a leave rule?
Leave rules are enforced by a policy engine written in code. The AI explains the engine's result but cannot overrule it, so a request that breaks a rule cannot be submitted.
03Does it work with our HRMS?
Yes — in a client deployment the assistant's tools connect to your HRMS (for example Keka, Darwinbox or Zoho People) and your leave policy is configured in the engine.
Need this built?

AI HR Manager

Your HR policies, answering employees 24/7.

Keep reading