Self-hosting an AI interviewer with LiveKit: the architecture
Real-time voice AI is a latency game. A look at how we combine self-hosted LiveKit, streaming speech and an LLM to run natural interviews on your own servers.
A text chatbot can take two seconds to reply and nobody minds. In a voice conversation, a two-second pause feels like the line dropped. Building an AI interviewer that candidates find natural is therefore mostly an exercise in shaving latency — while keeping sensitive audio inside infrastructure you control.
Why self-host LiveKit?
LiveKit is an open-source WebRTC media server. Hosting it yourself means candidate audio and video never transit a third-party media cloud, you choose the data-centre region, and you can satisfy privacy reviews that would block a SaaS tool. It also removes per-minute media fees at volume.
The pipeline
- Room: the candidate joins a LiveKit room from the browser — no install.
- Agent: a server-side agent joins the same room as a participant and subscribes to the candidate's audio track.
- Listen: streaming speech-to-text produces partial transcripts while the candidate is still talking; voice-activity detection decides when they have finished.
- Think: the LLM receives the transcript plus the role rubric and conversation state, and streams back the next question.
- Speak: streaming text-to-speech starts playing the first words before the full sentence is generated.
Where the milliseconds go
End-of-turn detection, first token from the model and first audio byte from speech synthesis dominate perceived latency. We tune turn detection per language, keep prompts compact and cache-friendly, and stream every stage so they overlap instead of queueing.
Making it fair and auditable
Every question is anchored to a structured rubric, every score cites the transcript moment that justifies it, and recordings are retained for human review. The AI screens and documents; people decide.
Running it in production
LiveKit, the agent workers and the application run behind Nginx with TURN for restrictive networks, horizontal agent scaling for interview spikes, and observability on turn latency so regressions are caught before candidates notice.
- #LiveKit
- #Voice AI
- #WebRTC
- #Architecture
AI Interviewer
Real-time voice interviews, run on your own servers.
Keep reading
ERP & Operations5 min read
From chatbot to co-worker: building an HR assistant that takes actions safely
AIVCJ HR doesn't just answer leave questions — it applies leave, routes approvals and drafts HR letters. Here is how we let an AI act without letting it make the rules: a policy engine in code, a confirmation card before every action, role permissions and an audit log. Measured on a scripted and a blind test set.
Read more2 min read
What is RAG — and when does your business actually need it?
Retrieval-augmented generation lets AI answer from your own documents instead of guessing. Here is how it works, where it shines, and the signs you are ready for it.
Read more
AI & RAG4 min read
Why most company chatbots hallucinate — and how we built one that cites its sources
A behind-the-scenes look at AIVCJ Knowledge: hybrid retrieval, query planning, a grounding check on every answer, and two public test sets — including a blind one written the way real people type. Try it live on sample HR, product and GST documents, or your own PDF.
Read more