LiveKit Agents
Capture available LiveKit Agents spans, milestones and enabled audio using a mixin and recorder, with coverage and shutdown limitations.
If your Python agent runs on LiveKit Agents, this integration records it with the
least code. It subscribes to AgentSession events and produces the same spans,
milestones and audio tracks you would otherwise write by hand.
Capture coverage depends on enabled settings, bound hooks, provider events and
successful writes. A frame observed in the agent pipeline is not proof of
downstream delivery or that a caller heard it. Review capture_status and the
limits below rather than assuming every call is complete.
Install
Check the build image before using git+https:// on LiveKit Cloud.
Public repository access does not supply a missing git binary or network
access in the build environment. A vendored wheel avoids those build-time
requirements — see
Deploying to LiveKit Cloud.
For local development, where you have git and network access to GitHub:
pip install "vaanieval-observer[livekit] @ git+https://github.com/shubhamofbce/vaanieval-observer-python-sdk.git"Source installs and hosted beta
The repositories are public and support anonymous Git access. These guides install the SDKs from public Git. Anonymous registry checks on 20 September 2026 returned HTTP 404 for npm @vaanieal/observer and PyPI vaanieval-observer; no release was found under those names. Source availability does not mean the hosted service is ready for public launch: new public workspaces remain closed.
The quickstarts pin published Python 0.5.7b1 and Node 0.1.1-beta.1 source previews. Both installed previews have uploaded synthetic calls to the isolated hosted development environment. Git/build tooling and explicit revision upgrades are required; this is not registry publication or production-readiness certification.
For setup help, email shubham@vaanieval.com or book a call.
Wire it up
from livekit.agents import Agent, AgentSession, WorkerOptions, cli
from vaani_observer.integrations.livekit import (
VaaniAudioTapMixin,
VaaniLiveKitRecorder,
observe_agent_session,
)
class MyAgent(VaaniAudioTapMixin, Agent):
def __init__(self) -> None:
super().__init__(instructions="You are a helpful support agent.")
async def entrypoint(ctx):
recorder = VaaniLiveKitRecorder.from_env(agent_id="support-bot")
agent = MyAgent()
session = AgentSession(...)
# Binds agent.vaani, subscribes to the session, and registers the shutdown
# hook that finalizes the recording when the job actually ends.
observe_agent_session(session, recorder, agent=agent, job_ctx=ctx)
await session.start(agent=agent, room=ctx.room)
if __name__ == "__main__":
cli.run_app(
WorkerOptions(
entrypoint_fnc=entrypoint,
# The default is 10 seconds, which is not enough to finalize and
# upload a call of any length. See "Uploading takes time".
shutdown_process_timeout=120.0,
)
)Never call finish() in a finally around session.start(). On
livekit-agents 1.x, start() returns as soon as the session starts, not when
the call ends. A finally therefore fires seconds into a live call: you get a
truncated recording, and every event after that point is discarded.
Passing job_ctx=ctx registers a shutdown callback instead, which runs when
the job genuinely ends. That is the only correct place to finalize.
VaaniAudioTapMixin must come before Agent in the base class list. Pass
agent= to observe_agent_session and it wires agent.vaani for you; if you
set it by hand and forget, you get spans and milestones but no audio and no
turn_id on instrumented LLM HTTP calls. The recorder warns when it is enabled
and no agent was bound.
Configure from the environment
from_env() mirrors the Node agent's variables so the same deployment config
works for both:
| Variable | Default | Meaning |
|---|---|---|
VAANI_ENABLED | false | Master switch. Off by default. |
VAANI_ENDPOINT | http://localhost:8000 | Dashboard base URL |
VAANI_API_KEY | local-dev | Bearer token for upload |
VAANI_SPOOL_DIR | .vaani-spool | Local package directory |
VAANI_CAPTURE_AUDIO | true | Record caller and agent PCM |
VAANI_CAPTURE_HTTP_BODIES | true | Capture request/response bodies |
VAANI_CAPTURE_STT_CONTENT | true | Capture transcript text |
VAANI_PAYLOAD_MAX_BYTES | 16384 | Payload size bound |
VAANI_AGENT_ID | livekit-agent | Default agent id |
VAANI_UPLOAD | true | Upload after finish() |
VAANI_ENDPOINTS | (unset) | JSON array of connection-capture rules |
If VAANI_ENABLED is true but recording cannot start, the recorder logs at
ERROR and recorder.last_error says why. It never raises: a broken
observability config must not be the reason a call fails to connect.
The integration's defaults are more permissive than the core SDK's:
VAANI_CAPTURE_HTTP_BODIES and VAANI_CAPTURE_STT_CONTENT default to true
here, so prompts and caller transcripts are captured in plain text unless you
set them to false. That is a deliberate choice for a debugging integration,
but it is almost certainly not what you want pointed at production traffic
without a privacy review. See
Capture and privacy.
Constructor options can be passed directly, overriding the environment:
recorder = VaaniLiveKitRecorder(
observer=None, # built from env if omitted
agent_id="support-bot",
metadata={"env": "prod", "version": "2026.08.1"},
capture_transcripts=True,
upload=False,
input_sample_rate=24000,
output_sample_rate=24000,
channels=1,
)from_env() never raises. If configuration is invalid or the SDK cannot
start, it returns an inert recorder whose methods are no-ops and whose
enabled property is False. Observability cannot take your agent down. The
cost is that a misconfiguration is silent — check recorder.enabled at startup
if you want to know.
What it records
STT spans, from the user's speech
user_state_changed and user_input_transcribed drive an STT span carrying
speech_started, first_partial, final_transcript and speech_final
milestones, plus partial transcripts as bounded samples. Once the sample limit is
reached it emits partial_samples_truncated rather than growing without bound.
A turn ends where LiveKit ends it, which is the commit — not the final
transcript. A provider is free to end a transcript at every sentence, and
LiveKit merges those segments into one user message and answers it once; it
also answers a provisional end of turn, creating a reply for each final and
cancelling the ones the caller talks past. Segments that belong to one
committed message therefore share a single STT span, which reports how many
arrived in response.final_segments, and the reply that survives is recorded
against that same turn. Recorded turns match the framework's own conversation
history where the integration can observe the required events. Ignored
utterances, split turns and inferred attribution need the caveats below.
Utterances your agent ignores
An agent can decline to answer by raising StopResponse from
Agent.on_user_turn_completed — a content filter, a wake-word gate, a barge-in
policy. LiveKit ends that turn and clears its transcript, and no session event
announces it. Left alone, the next thing the caller says looks like a
continuation of the ignored utterance and is recorded prepended to it, so one
turn ends up carrying two questions.
When you pass agent= to observe_agent_session, the SDK wraps that method to
learn about it. The ignored utterance stays its own turn, keeps its transcript,
and its STT span carries response.reply_skipped: "stop_response". The wrapper
is transparent: your return value and your exceptions pass through untouched.
Note that a turn you decline can still cost money. LiveKit may have started a preemptive LLM call before your filter ran, and that request is billed whether or not it is spoken. Those tokens are recorded against the turn whose audio caused them, so the dashboard shows an ignored turn with a real model bill and no speech out, labelled as declined rather than flagged as a silent failure.
Any other exception from on_user_turn_completed is caught by LiveKit, logged
and swallowed: the message is never committed and no reply is generated. That
records as reply_skipped: "callback_error", so an agent whose turn callback is
failing is not mistaken for one that has gone mute — the two need opposite
responses and the numbers alone cannot tell them apart.
If the agent object cannot be wrapped (a frozen or slotted subclass), the SDK warns once and every committed turn is still recorded correctly — only the ignored-utterance case becomes unobservable.
If you hand off to another agent mid-call with AgentSession.update_agent(),
the SDK follows the handoff and instruments the replacement. Without that the
replacement carries neither the audio tap nor the turn watch, and the rest of
the call records no agent audio at all.
When one message is recorded as two turns
A recogniser can deliver a single utterance as several finals, and LiveKit
merges them into one user message. We normally merge them too. There is one
case where we cannot: if something has already closed and published the span
holding the earlier words — a filler say(), or a preemptive reply that was
billed against the unfinished message — then merging would write into a
published span and the caller's words would be lost silently.
We keep the words and open a second turn instead. That turn carries
continues_turn naming the first half, and split_reason. The dashboard
labels it rather than showing two unrelated turns, because the first half
otherwise reads as a caller talking to an agent that never answered, and
per-turn averages would be computed over halves.
A say() spoken while the caller's message is still open is treated as a
filler, not as its answer, and is recorded as its own turn. Once LiveKit has
ended that message a say() is the opposite case — the scripted answer of an
agent that then raises StopResponse — and is recorded against the question it
answers.
LLM spans, from metrics
metrics_collected produces an LLM span with a first_token milestone placed at
started_at + ttft, and token counts when the provider reports them.
If your agent's LLM emits no llm_metrics at all — a custom llm_node, or a
provider plugin without a metrics implementation — the span is instead derived
from conversation_item_added. A derived span carries
request.derived_from: "conversation_item_added" and response.estimated: true,
has approximate timings, and reports no token counts rather than inventing
zeros. The recorder warns once when it falls back.
Check the model label. openai.LLM.with_azure() defaults its model
argument to "gpt-4o" regardless of which deployment you point it at, so the
plugin reports gpt-4o for a gpt-5-mini deployment. The integration detects
the real Azure deployment name and records that, keeping the plugin's claim as
request.reported_model. If the label is still wrong, set it explicitly:
VaaniLiveKitRecorder.from_env(model_overrides={"llm": "gpt-5-mini"})TTS spans, with real audio accounting
A speak milestone with the character count, a first_byte milestone at
started_at + ttfb, and a response carrying three separate numbers. A reply
that was cut off is ended with status cancelled, not error — barge-in is
correct behaviour.
| Field | Source | Means |
|---|---|---|
audio_ms | the TTS provider's own metric | how much audio it synthesized |
played_ms | the PCM captured in tts_node | duration observed in the agent output pipeline, not independently verified caller delivery |
audio_bytes | the PCM captured in tts_node | the raw byte count behind played_ms |
On a cancelled span played_ms is normally lower than audio_ms, and
that gap is useful for reviewing interrupted output. It does not independently
measure what the caller heard. They come from different sources, so treat them as two measurements
rather than one number reported twice. played_ms is omitted entirely — not
reported as 0 — when no audio was captured, so missing capture is not
confused with measured silence.
Tool calls
function_tools_executed produces tool operations on the owning turn.
Audio, from the node hooks
stt_node and tts_node are LiveKit's supported extension points, so the mixin
tees the exact frames the pipeline uses rather than reaching into private io
plumbing that changes between releases.
End-of-utterance timing
_record_eou attaches LiveKit's own end-of-utterance measurement, which is the
best available source for endpointing latency.
Errors are routed to the component that failed
A session error closes the spans of the component that actually failed, rather than marking the whole turn bad:
| LiveKit error | Spans closed as failed |
|---|---|
stt_error | STT |
llm_error | LLM |
tts_error | TTS |
realtime_model_error | STT, LLM and TTS |
recorder.fail(error) records that the call could not be run at all.
Version drift
The integration subscribes to nine AgentSession events, each guarded
individually — a handler that raises is logged and swallowed, because an
exception on LiveKit's event loop would kill the call.
metrics_collected is deprecated in LiveKit 1.6 in favour of
session_usage_updated plus ChatMessage.metrics, but it remains the only
source of per-stage duration, TTFT/TTFB and token counts. The integration
subscribes to both, so spans survive its removal with only the token counts
degrading. Subscription failures are logged at debug level and skipped.
Manual escape hatches
The recorder exposes the underlying primitives when the automatic path is not enough:
| Method | Purpose |
|---|---|
recorder.call | The underlying Session (None when inert) |
recorder.enabled | Whether recording is active |
recorder.attach(session) | Subscribe to another AgentSession |
recorder.turn_context() | Scope ambient work to the turn being served |
recorder.observe_socket(socket, url=..., endpoint_id=...) | Record a provider socket |
recorder.tap_input_frame(frame) / tap_output_frame(frame) | Feed PCM manually |
recorder.finalize_open_spans(outcome=...) | Close everything still open |
recorder.fail(error) | Record a call that could not run |
await recorder.finish(outcome=...) | Finalize and, if configured, upload |
Constants for the endpoint ids the integration uses:
from vaani_observer.integrations.livekit import (
STT_ENDPOINT_ID, # "stt"
LLM_ENDPOINT_ID, # "llm"
TTS_ENDPOINT_ID, # "tts"
)Uploading takes time
This is the part that most often goes wrong in production, so it is worth understanding before you deploy.
Audio is recorded as raw pcm_s16le. At 24 kHz stereo that is 94 KB/s, so a
five-minute call is about 28 MB. The SDK gzips it in transit, but how much
that helps depends entirely on the audio:
| Content | Measured ratio |
|---|---|
| Real call audio (speech with silence) | ~38% of the original |
| Dense or noisy audio | as little as 93% — effectively no gain |
Plan for the pessimistic case. On a 10 Mbps uplink, 28 MB is still over twenty seconds, and a longer call is proportionally worse.
LiveKit's default shutdown_process_timeout is 10 seconds. If the upload has
not finished by then, the process is killed mid-transfer.
| Setting | Why |
|---|---|
shutdown_process_timeout=120.0 | Gives the upload room to finish |
VAANI_SPOOL_DIR on a persistent volume | So a killed upload can be retried |
python -m vaani_observer.drain | Ships whatever was left behind |
On LiveKit Cloud the spool is ephemeral. It disappears with the container, so a package that misses the shutdown budget there is gone for good. Either raise the timeout enough to cover your longest call, or run the drainer as a sidecar sharing a volume with the worker.
If the upload fails or runs out of budget, the recorder normally keeps the package and logs its directory. A hard kill or ephemeral-volume teardown can still lose data; persistent storage is needed for later retry:
# One-shot, after a batch of calls
python -m vaani_observer.drain --verbose
# Or as a sidecar, re-scanning every 60 seconds
python -m vaani_observer.drain --watch 60The drainer reads VAANI_ENDPOINT, VAANI_API_KEY and VAANI_SPOOL_DIR, uploads
every finalized package it finds, and removes each one once the backend has
verified its digests. Pass --keep to leave them on disk instead.
Each package gets --timeout seconds (default 1800) to ship. A pass is
sequential, so without a budget one stalled upload holds up every recording
queued behind it. Pass --timeout 0 to remove the bound.
Receipts are scoped to the endpoint that issued them. If you repoint
VAANI_ENDPOINT — staging to production, self-hosted to cloud — packages
already delivered elsewhere are treated as pending again and re-uploaded to the
new destination, which is idempotent. A package is only ever removed when its
receipt names the backend you are currently writing to. A receipt that names no
backend — one written by an older version — still suppresses a re-upload, but is
never treated as grounds for deleting the call.
Repointing requires --yes. A spool records every endpoint it has been
used with, in .vaani-destinations at its root. A drain configured for a
destination that file has never seen refuses the whole pass, exits 2 and
uploads nothing:
refusing to drain: /srv/spool has previously been used with
https://ingest.example.com but this pass is configured for
https://ingest.exmaple.com. 5 package(s) hold raw call audio and caller
speech; uploading them would disclose that to a host this spool has never
shipped to. Confirm with --yes if the change is intended, or correct
VAANI_ENDPOINT.Packages contain the caller's voice and whatever they said, and a successful upload is what authorises deleting the local copy — so a mistyped hostname would otherwise disclose the recordings and destroy the only other copy, reporting a clean pass. A confirmed change uploads but keeps the local copies; the next pass, once the new endpoint is a known one, reclaims the disk normally.
It also removes packages that were uploaded in-process by finish() — those
carry an uploaded.json receipt and are otherwise skipped by every pass, which
at ~28 MB per five-minute call fills a busy worker's disk. They are kept for
--delivered-ttl hours (default 24) first, because during an incident the
local package is the only evidence you have. Pass a negative value to keep them
indefinitely.
Removal relies on the backend receipt: it proves digest verification at upload
time, not continuing availability, backup or retention of the remote copy.
Choose the cleanup window accordingly. Age comes from the uploaded_at
field inside the receipt rather than the file's mtime, which any copy resets.
Upload tuning
recorder = VaaniLiveKitRecorder.from_env()
observe_agent_session(session, recorder, agent=agent, job_ctx=ctx,
upload_timeout=90.0) # budget for the whole uploadfrom_env() builds the observer for you, so these are set through the
environment:
| Variable | Default | Meaning |
|---|---|---|
VAANI_UPLOAD_TIMEOUT_S | 30.0 | Socket timeout for the handshake, before the size allowance below |
VAANI_UPLOAD_RETRIES | 3 | Retries on 5xx, 429 and dropped connections |
VAANI_UPLOAD_COMPRESS | 1 | gzip objects of 64 KiB or more in transit |
VAANI_UPLOAD_MIN_THROUGHPUT_BPS | 131072 | Assumed worst-case upload speed |
Large recordings need a transfer budget. The SDK extends the base request
timeout using the object size and VAANI_UPLOAD_MIN_THROUGHPUT_BPS. Lowering
the assumed throughput permits longer attempts; raising it fails slow
transfers sooner. The overall upload budget and the host's shutdown deadline
still limit available time. Test your longest calls on the actual uplink and
keep a persistent spool rather than treating retries as guaranteed delivery.
Hard ceiling: 128 MiB per object, which is about 23 minutes of 24 kHz stereo audio in the local API. A longer object is rejected at upload time; keep the local package on persistent storage for recovery. For future calls, record at a lower sample rate or turn audio capture off and keep the spans.
Deploying to LiveKit Cloud
The Git repository is public, but a git+https:// install still
requires git and network access in your build image. If either is unavailable,
build a wheel from an approved revision and vendor it instead:
# From a checkout of the SDK
python -m build --wheel
cp dist/vaanieval_observer-*.whl /path/to/your/agent/vendor/livekit-agents[openai,deepgram,silero]~=1.6
./vendor/vaanieval_observer-0.5.6-py3-none-any.whl[livekit]Vendoring makes the build less dependent on GitHub availability, but you must refresh the wheel deliberately for fixes and keep its source revision recorded. Use the filename produced by your build rather than assuming the example version is a current release.
Then set VAANI_ENABLED=true, VAANI_ENDPOINT and VAANI_API_KEY in your
LiveKit Cloud environment, and raise shutdown_process_timeout as above.
This describes deploying your agent, not opening the local dashboard to the
internet. app.main:app must remain localhost-only; a LiveKit Cloud worker
cannot reach your laptop's localhost. Use local capture with VAANI_UPLOAD=false
until you have an approved compatible destination and a persistent-spool plan.
The separate hosted-beta backend is under construction, not a generally
available upload endpoint.
Keep the spool out of your repo and your images
The spool holds raw call audio and transcripts — the most sensitive data
your agent touches. It defaults to .vaani-spool/ in the working directory,
which is usually your project root, so add it to both ignore files before the
first recording:
.vaani-spool/.vaani-spool/Without the .dockerignore entry, any packages left on the spool from a local
test run are copied into your image by COPY . . and ship to production —
and to anyone who can pull the image. Without the .gitignore entry they are
one git add . away from your history.
Capturing provider connections
By default the integration records spans from LiveKit's own metrics and does
not patch httpx or aiohttp. Point VAANI_ENDPOINTS at an endpoint to
capture its connections too:
VAANI_ENDPOINTS='[{"id":"deepgram","type":"stt","url":"wss://api.deepgram.com","match":"origin"}]'Do not add a rule of type llm for the endpoint your agent's LLM plugin
already uses. You would get two LLM spans per turn — one from
metrics_collected and one from the HTTP capture — and every token and latency
aggregate built on the recording would double-count. Use this for endpoints
LiveKit metrics do not already describe.
Limits worth knowing
- Only
AgentSessionis supported. A custom pipeline built directly on LiveKit primitives needs manual instrumentation. - Span quality depends on what the provider reports. TTFT and TTFB come from LiveKit metrics; a provider plugin that does not report them yields a span with duration but no milestones, and those turns count as unmeasurable.
- Finalization happens in the job's shutdown hook. A
SIGKILL— an OOM, or a shutdown budget too small for the upload — still loses whatever had not reached disk. Files that reached a persistent spool may be recoverable; ephemeral container storage is not a durability guarantee. - The recorder is per-session. Reusing one across concurrent
AgentSessions will interleave turns into one recording. - A degraded recording says so. If audio is dropped mid-call, the manifest
reports
audio_complete: falserather than claiming a complete recording over a gap.