Skip to content
Engineering

Self-hosting an AI interviewer with LiveKit: the architecture

Real-time voice AI is a latency game. A look at how we combine self-hosted LiveKit, streaming speech and an LLM to run natural interviews on your own servers.

1 min read

A text chatbot can take two seconds to reply and nobody minds. In a voice conversation, a two-second pause feels like the line dropped. Building an AI interviewer that candidates find natural is therefore mostly an exercise in shaving latency — while keeping sensitive audio inside infrastructure you control.

Why self-host LiveKit?

LiveKit is an open-source WebRTC media server. Hosting it yourself means candidate audio and video never transit a third-party media cloud, you choose the data-centre region, and you can satisfy privacy reviews that would block a SaaS tool. It also removes per-minute media fees at volume.

The pipeline

  1. Room: the candidate joins a LiveKit room from the browser — no install.
  2. Agent: a server-side agent joins the same room as a participant and subscribes to the candidate's audio track.
  3. Listen: streaming speech-to-text produces partial transcripts while the candidate is still talking; voice-activity detection decides when they have finished.
  4. Think: the LLM receives the transcript plus the role rubric and conversation state, and streams back the next question.
  5. Speak: streaming text-to-speech starts playing the first words before the full sentence is generated.

Where the milliseconds go

End-of-turn detection, first token from the model and first audio byte from speech synthesis dominate perceived latency. We tune turn detection per language, keep prompts compact and cache-friendly, and stream every stage so they overlap instead of queueing.

Making it fair and auditable

Every question is anchored to a structured rubric, every score cites the transcript moment that justifies it, and recordings are retained for human review. The AI screens and documents; people decide.

Running it in production

LiveKit, the agent workers and the application run behind Nginx with TURN for restrictive networks, horizontal agent scaling for interview spikes, and observability on turn latency so regressions are caught before candidates notice.

  • #LiveKit
  • #Voice AI
  • #WebRTC
  • #Architecture
LinkedInWhatsApp
Need this built?

AI Interviewer

Real-time voice interviews, run on your own servers.

Keep reading