Vaanieval
Python SDK

LiveKit Agents

Capture available LiveKit Agents spans, milestones and enabled audio using a mixin and recorder, with coverage and shutdown limitations.

If your Python agent runs on LiveKit Agents, this integration records it with the least code. It subscribes to AgentSession events and produces the same spans, milestones and audio tracks you would otherwise write by hand.

Capture coverage depends on enabled settings, bound hooks, provider events and successful writes. A frame observed in the agent pipeline is not proof of downstream delivery or that a caller heard it. Review capture_status and the limits below rather than assuming every call is complete.

Install

Check the build image before using git+https:// on LiveKit Cloud. Public repository access does not supply a missing git binary or network access in the build environment. A vendored wheel avoids those build-time requirements — see Deploying to LiveKit Cloud.

For local development, where you have git and network access to GitHub:

pip install "vaanieval-observer[livekit] @ git+https://github.com/shubhamofbce/vaanieval-observer-python-sdk.git"

Source installs and hosted beta

The repositories are public and support anonymous Git access. These guides install the SDKs from public Git. Anonymous registry checks on 20 September 2026 returned HTTP 404 for npm @vaanieal/observer and PyPI vaanieval-observer; no release was found under those names. Source availability does not mean the hosted service is ready for public launch: new public workspaces remain closed.

The quickstarts pin published Python 0.5.7b1 and Node 0.1.1-beta.1 source previews. Both installed previews have uploaded synthetic calls to the isolated hosted development environment. Git/build tooling and explicit revision upgrades are required; this is not registry publication or production-readiness certification.

For setup help, email shubham@vaanieval.com or book a call.

Wire it up

agent.py
from livekit.agents import Agent, AgentSession, WorkerOptions, cli
from vaani_observer.integrations.livekit import (
    VaaniAudioTapMixin,
    VaaniLiveKitRecorder,
    observe_agent_session,
)


class MyAgent(VaaniAudioTapMixin, Agent):
    def __init__(self) -> None:
        super().__init__(instructions="You are a helpful support agent.")


async def entrypoint(ctx):
    recorder = VaaniLiveKitRecorder.from_env(agent_id="support-bot")

    agent = MyAgent()
    session = AgentSession(...)

    # Binds agent.vaani, subscribes to the session, and registers the shutdown
    # hook that finalizes the recording when the job actually ends.
    observe_agent_session(session, recorder, agent=agent, job_ctx=ctx)

    await session.start(agent=agent, room=ctx.room)


if __name__ == "__main__":
    cli.run_app(
        WorkerOptions(
            entrypoint_fnc=entrypoint,
            # The default is 10 seconds, which is not enough to finalize and
            # upload a call of any length. See "Uploading takes time".
            shutdown_process_timeout=120.0,
        )
    )

Never call finish() in a finally around session.start(). On livekit-agents 1.x, start() returns as soon as the session starts, not when the call ends. A finally therefore fires seconds into a live call: you get a truncated recording, and every event after that point is discarded.

Passing job_ctx=ctx registers a shutdown callback instead, which runs when the job genuinely ends. That is the only correct place to finalize.

VaaniAudioTapMixin must come before Agent in the base class list. Pass agent= to observe_agent_session and it wires agent.vaani for you; if you set it by hand and forget, you get spans and milestones but no audio and no turn_id on instrumented LLM HTTP calls. The recorder warns when it is enabled and no agent was bound.

Configure from the environment

from_env() mirrors the Node agent's variables so the same deployment config works for both:

VariableDefaultMeaning
VAANI_ENABLEDfalseMaster switch. Off by default.
VAANI_ENDPOINThttp://localhost:8000Dashboard base URL
VAANI_API_KEYlocal-devBearer token for upload
VAANI_SPOOL_DIR.vaani-spoolLocal package directory
VAANI_CAPTURE_AUDIOtrueRecord caller and agent PCM
VAANI_CAPTURE_HTTP_BODIEStrueCapture request/response bodies
VAANI_CAPTURE_STT_CONTENTtrueCapture transcript text
VAANI_PAYLOAD_MAX_BYTES16384Payload size bound
VAANI_AGENT_IDlivekit-agentDefault agent id
VAANI_UPLOADtrueUpload after finish()
VAANI_ENDPOINTS(unset)JSON array of connection-capture rules

If VAANI_ENABLED is true but recording cannot start, the recorder logs at ERROR and recorder.last_error says why. It never raises: a broken observability config must not be the reason a call fails to connect.

The integration's defaults are more permissive than the core SDK's: VAANI_CAPTURE_HTTP_BODIES and VAANI_CAPTURE_STT_CONTENT default to true here, so prompts and caller transcripts are captured in plain text unless you set them to false. That is a deliberate choice for a debugging integration, but it is almost certainly not what you want pointed at production traffic without a privacy review. See Capture and privacy.

Constructor options can be passed directly, overriding the environment:

recorder = VaaniLiveKitRecorder(
    observer=None,               # built from env if omitted
    agent_id="support-bot",
    metadata={"env": "prod", "version": "2026.08.1"},
    capture_transcripts=True,
    upload=False,
    input_sample_rate=24000,
    output_sample_rate=24000,
    channels=1,
)

from_env() never raises. If configuration is invalid or the SDK cannot start, it returns an inert recorder whose methods are no-ops and whose enabled property is False. Observability cannot take your agent down. The cost is that a misconfiguration is silent — check recorder.enabled at startup if you want to know.

What it records

STT spans, from the user's speech

user_state_changed and user_input_transcribed drive an STT span carrying speech_started, first_partial, final_transcript and speech_final milestones, plus partial transcripts as bounded samples. Once the sample limit is reached it emits partial_samples_truncated rather than growing without bound.

A turn ends where LiveKit ends it, which is the commit — not the final transcript. A provider is free to end a transcript at every sentence, and LiveKit merges those segments into one user message and answers it once; it also answers a provisional end of turn, creating a reply for each final and cancelling the ones the caller talks past. Segments that belong to one committed message therefore share a single STT span, which reports how many arrived in response.final_segments, and the reply that survives is recorded against that same turn. Recorded turns match the framework's own conversation history where the integration can observe the required events. Ignored utterances, split turns and inferred attribution need the caveats below.

Utterances your agent ignores

An agent can decline to answer by raising StopResponse from Agent.on_user_turn_completed — a content filter, a wake-word gate, a barge-in policy. LiveKit ends that turn and clears its transcript, and no session event announces it. Left alone, the next thing the caller says looks like a continuation of the ignored utterance and is recorded prepended to it, so one turn ends up carrying two questions.

When you pass agent= to observe_agent_session, the SDK wraps that method to learn about it. The ignored utterance stays its own turn, keeps its transcript, and its STT span carries response.reply_skipped: "stop_response". The wrapper is transparent: your return value and your exceptions pass through untouched.

Note that a turn you decline can still cost money. LiveKit may have started a preemptive LLM call before your filter ran, and that request is billed whether or not it is spoken. Those tokens are recorded against the turn whose audio caused them, so the dashboard shows an ignored turn with a real model bill and no speech out, labelled as declined rather than flagged as a silent failure.

Any other exception from on_user_turn_completed is caught by LiveKit, logged and swallowed: the message is never committed and no reply is generated. That records as reply_skipped: "callback_error", so an agent whose turn callback is failing is not mistaken for one that has gone mute — the two need opposite responses and the numbers alone cannot tell them apart.

If the agent object cannot be wrapped (a frozen or slotted subclass), the SDK warns once and every committed turn is still recorded correctly — only the ignored-utterance case becomes unobservable.

If you hand off to another agent mid-call with AgentSession.update_agent(), the SDK follows the handoff and instruments the replacement. Without that the replacement carries neither the audio tap nor the turn watch, and the rest of the call records no agent audio at all.

When one message is recorded as two turns

A recogniser can deliver a single utterance as several finals, and LiveKit merges them into one user message. We normally merge them too. There is one case where we cannot: if something has already closed and published the span holding the earlier words — a filler say(), or a preemptive reply that was billed against the unfinished message — then merging would write into a published span and the caller's words would be lost silently.

We keep the words and open a second turn instead. That turn carries continues_turn naming the first half, and split_reason. The dashboard labels it rather than showing two unrelated turns, because the first half otherwise reads as a caller talking to an agent that never answered, and per-turn averages would be computed over halves.

A say() spoken while the caller's message is still open is treated as a filler, not as its answer, and is recorded as its own turn. Once LiveKit has ended that message a say() is the opposite case — the scripted answer of an agent that then raises StopResponse — and is recorded against the question it answers.

LLM spans, from metrics

metrics_collected produces an LLM span with a first_token milestone placed at started_at + ttft, and token counts when the provider reports them.

If your agent's LLM emits no llm_metrics at all — a custom llm_node, or a provider plugin without a metrics implementation — the span is instead derived from conversation_item_added. A derived span carries request.derived_from: "conversation_item_added" and response.estimated: true, has approximate timings, and reports no token counts rather than inventing zeros. The recorder warns once when it falls back.

Check the model label. openai.LLM.with_azure() defaults its model argument to "gpt-4o" regardless of which deployment you point it at, so the plugin reports gpt-4o for a gpt-5-mini deployment. The integration detects the real Azure deployment name and records that, keeping the plugin's claim as request.reported_model. If the label is still wrong, set it explicitly:

VaaniLiveKitRecorder.from_env(model_overrides={"llm": "gpt-5-mini"})

TTS spans, with real audio accounting

A speak milestone with the character count, a first_byte milestone at started_at + ttfb, and a response carrying three separate numbers. A reply that was cut off is ended with status cancelled, not error — barge-in is correct behaviour.

FieldSourceMeans
audio_msthe TTS provider's own metrichow much audio it synthesized
played_msthe PCM captured in tts_nodeduration observed in the agent output pipeline, not independently verified caller delivery
audio_bytesthe PCM captured in tts_nodethe raw byte count behind played_ms

On a cancelled span played_ms is normally lower than audio_ms, and that gap is useful for reviewing interrupted output. It does not independently measure what the caller heard. They come from different sources, so treat them as two measurements rather than one number reported twice. played_ms is omitted entirely — not reported as 0 — when no audio was captured, so missing capture is not confused with measured silence.

Tool calls

function_tools_executed produces tool operations on the owning turn.

Audio, from the node hooks

stt_node and tts_node are LiveKit's supported extension points, so the mixin tees the exact frames the pipeline uses rather than reaching into private io plumbing that changes between releases.

End-of-utterance timing

_record_eou attaches LiveKit's own end-of-utterance measurement, which is the best available source for endpointing latency.

Errors are routed to the component that failed

A session error closes the spans of the component that actually failed, rather than marking the whole turn bad:

LiveKit errorSpans closed as failed
stt_errorSTT
llm_errorLLM
tts_errorTTS
realtime_model_errorSTT, LLM and TTS

recorder.fail(error) records that the call could not be run at all.

Version drift

The integration subscribes to nine AgentSession events, each guarded individually — a handler that raises is logged and swallowed, because an exception on LiveKit's event loop would kill the call.

metrics_collected is deprecated in LiveKit 1.6 in favour of session_usage_updated plus ChatMessage.metrics, but it remains the only source of per-stage duration, TTFT/TTFB and token counts. The integration subscribes to both, so spans survive its removal with only the token counts degrading. Subscription failures are logged at debug level and skipped.

Manual escape hatches

The recorder exposes the underlying primitives when the automatic path is not enough:

MethodPurpose
recorder.callThe underlying Session (None when inert)
recorder.enabledWhether recording is active
recorder.attach(session)Subscribe to another AgentSession
recorder.turn_context()Scope ambient work to the turn being served
recorder.observe_socket(socket, url=..., endpoint_id=...)Record a provider socket
recorder.tap_input_frame(frame) / tap_output_frame(frame)Feed PCM manually
recorder.finalize_open_spans(outcome=...)Close everything still open
recorder.fail(error)Record a call that could not run
await recorder.finish(outcome=...)Finalize and, if configured, upload

Constants for the endpoint ids the integration uses:

from vaani_observer.integrations.livekit import (
    STT_ENDPOINT_ID,   # "stt"
    LLM_ENDPOINT_ID,   # "llm"
    TTS_ENDPOINT_ID,   # "tts"
)

Uploading takes time

This is the part that most often goes wrong in production, so it is worth understanding before you deploy.

Audio is recorded as raw pcm_s16le. At 24 kHz stereo that is 94 KB/s, so a five-minute call is about 28 MB. The SDK gzips it in transit, but how much that helps depends entirely on the audio:

ContentMeasured ratio
Real call audio (speech with silence)~38% of the original
Dense or noisy audioas little as 93% — effectively no gain

Plan for the pessimistic case. On a 10 Mbps uplink, 28 MB is still over twenty seconds, and a longer call is proportionally worse.

LiveKit's default shutdown_process_timeout is 10 seconds. If the upload has not finished by then, the process is killed mid-transfer.

SettingWhy
shutdown_process_timeout=120.0Gives the upload room to finish
VAANI_SPOOL_DIR on a persistent volumeSo a killed upload can be retried
python -m vaani_observer.drainShips whatever was left behind

On LiveKit Cloud the spool is ephemeral. It disappears with the container, so a package that misses the shutdown budget there is gone for good. Either raise the timeout enough to cover your longest call, or run the drainer as a sidecar sharing a volume with the worker.

If the upload fails or runs out of budget, the recorder normally keeps the package and logs its directory. A hard kill or ephemeral-volume teardown can still lose data; persistent storage is needed for later retry:

# One-shot, after a batch of calls
python -m vaani_observer.drain --verbose

# Or as a sidecar, re-scanning every 60 seconds
python -m vaani_observer.drain --watch 60

The drainer reads VAANI_ENDPOINT, VAANI_API_KEY and VAANI_SPOOL_DIR, uploads every finalized package it finds, and removes each one once the backend has verified its digests. Pass --keep to leave them on disk instead.

Each package gets --timeout seconds (default 1800) to ship. A pass is sequential, so without a budget one stalled upload holds up every recording queued behind it. Pass --timeout 0 to remove the bound.

Receipts are scoped to the endpoint that issued them. If you repoint VAANI_ENDPOINT — staging to production, self-hosted to cloud — packages already delivered elsewhere are treated as pending again and re-uploaded to the new destination, which is idempotent. A package is only ever removed when its receipt names the backend you are currently writing to. A receipt that names no backend — one written by an older version — still suppresses a re-upload, but is never treated as grounds for deleting the call.

Repointing requires --yes. A spool records every endpoint it has been used with, in .vaani-destinations at its root. A drain configured for a destination that file has never seen refuses the whole pass, exits 2 and uploads nothing:

refusing to drain: /srv/spool has previously been used with
https://ingest.example.com but this pass is configured for
https://ingest.exmaple.com. 5 package(s) hold raw call audio and caller
speech; uploading them would disclose that to a host this spool has never
shipped to. Confirm with --yes if the change is intended, or correct
VAANI_ENDPOINT.

Packages contain the caller's voice and whatever they said, and a successful upload is what authorises deleting the local copy — so a mistyped hostname would otherwise disclose the recordings and destroy the only other copy, reporting a clean pass. A confirmed change uploads but keeps the local copies; the next pass, once the new endpoint is a known one, reclaims the disk normally.

It also removes packages that were uploaded in-process by finish() — those carry an uploaded.json receipt and are otherwise skipped by every pass, which at ~28 MB per five-minute call fills a busy worker's disk. They are kept for --delivered-ttl hours (default 24) first, because during an incident the local package is the only evidence you have. Pass a negative value to keep them indefinitely.

Removal relies on the backend receipt: it proves digest verification at upload time, not continuing availability, backup or retention of the remote copy. Choose the cleanup window accordingly. Age comes from the uploaded_at field inside the receipt rather than the file's mtime, which any copy resets.

Upload tuning

recorder = VaaniLiveKitRecorder.from_env()
observe_agent_session(session, recorder, agent=agent, job_ctx=ctx,
                      upload_timeout=90.0)  # budget for the whole upload

from_env() builds the observer for you, so these are set through the environment:

VariableDefaultMeaning
VAANI_UPLOAD_TIMEOUT_S30.0Socket timeout for the handshake, before the size allowance below
VAANI_UPLOAD_RETRIES3Retries on 5xx, 429 and dropped connections
VAANI_UPLOAD_COMPRESS1gzip objects of 64 KiB or more in transit
VAANI_UPLOAD_MIN_THROUGHPUT_BPS131072Assumed worst-case upload speed

Large recordings need a transfer budget. The SDK extends the base request timeout using the object size and VAANI_UPLOAD_MIN_THROUGHPUT_BPS. Lowering the assumed throughput permits longer attempts; raising it fails slow transfers sooner. The overall upload budget and the host's shutdown deadline still limit available time. Test your longest calls on the actual uplink and keep a persistent spool rather than treating retries as guaranteed delivery.

Hard ceiling: 128 MiB per object, which is about 23 minutes of 24 kHz stereo audio in the local API. A longer object is rejected at upload time; keep the local package on persistent storage for recovery. For future calls, record at a lower sample rate or turn audio capture off and keep the spans.

Deploying to LiveKit Cloud

The Git repository is public, but a git+https:// install still requires git and network access in your build image. If either is unavailable, build a wheel from an approved revision and vendor it instead:

# From a checkout of the SDK
python -m build --wheel
cp dist/vaanieval_observer-*.whl /path/to/your/agent/vendor/
requirements.txt
livekit-agents[openai,deepgram,silero]~=1.6
./vendor/vaanieval_observer-0.5.6-py3-none-any.whl[livekit]

Vendoring makes the build less dependent on GitHub availability, but you must refresh the wheel deliberately for fixes and keep its source revision recorded. Use the filename produced by your build rather than assuming the example version is a current release.

Then set VAANI_ENABLED=true, VAANI_ENDPOINT and VAANI_API_KEY in your LiveKit Cloud environment, and raise shutdown_process_timeout as above.

This describes deploying your agent, not opening the local dashboard to the internet. app.main:app must remain localhost-only; a LiveKit Cloud worker cannot reach your laptop's localhost. Use local capture with VAANI_UPLOAD=false until you have an approved compatible destination and a persistent-spool plan. The separate hosted-beta backend is under construction, not a generally available upload endpoint.

Keep the spool out of your repo and your images

The spool holds raw call audio and transcripts — the most sensitive data your agent touches. It defaults to .vaani-spool/ in the working directory, which is usually your project root, so add it to both ignore files before the first recording:

.gitignore
.vaani-spool/
.dockerignore
.vaani-spool/

Without the .dockerignore entry, any packages left on the spool from a local test run are copied into your image by COPY . . and ship to production — and to anyone who can pull the image. Without the .gitignore entry they are one git add . away from your history.

Capturing provider connections

By default the integration records spans from LiveKit's own metrics and does not patch httpx or aiohttp. Point VAANI_ENDPOINTS at an endpoint to capture its connections too:

VAANI_ENDPOINTS='[{"id":"deepgram","type":"stt","url":"wss://api.deepgram.com","match":"origin"}]'

Do not add a rule of type llm for the endpoint your agent's LLM plugin already uses. You would get two LLM spans per turn — one from metrics_collected and one from the HTTP capture — and every token and latency aggregate built on the recording would double-count. Use this for endpoints LiveKit metrics do not already describe.

Limits worth knowing

  • Only AgentSession is supported. A custom pipeline built directly on LiveKit primitives needs manual instrumentation.
  • Span quality depends on what the provider reports. TTFT and TTFB come from LiveKit metrics; a provider plugin that does not report them yields a span with duration but no milestones, and those turns count as unmeasurable.
  • Finalization happens in the job's shutdown hook. A SIGKILL — an OOM, or a shutdown budget too small for the upload — still loses whatever had not reached disk. Files that reached a persistent spool may be recoverable; ephemeral container storage is not a durability guarantee.
  • The recorder is per-session. Reusing one across concurrent AgentSessions will interleave turns into one recording.
  • A degraded recording says so. If audio is dropped mid-call, the manifest reports audio_complete: false rather than claiming a complete recording over a gap.

Next

On this page