Troubleshooting
Symptoms you will actually hit, what causes them, and how to fix them — from empty traces to failed uploads to missing metrics.
Nothing is recorded at all
Check, in order:
- Was the observer constructed? Instrumentation is installed in the constructor. A lazily created observer that never runs records nothing.
- Did
session.end()run? The package directory exists during the call, butmanifest.jsonis only written byend(). Put it in afinallyblock. With the LiveKit integration, passjob_ctx=ctxtoobserve_agent_session()instead — afinallythere fires mid-call. - Is
spool_directorywritable? Withstrict=Falsea write failure is swallowed. Setstrict=Trueto see it. - LiveKit only:
VAANI_ENABLEDdefaults tofalse.from_env()returns an inert recorder when disabled. Checkrecorder.enabled, and readrecorder.last_error— whenVAANI_ENABLEDis set but recording is off, the reason is also logged at ERROR. - Did the upload run out of time? The package is kept on the spool and the
directory is logged. Ship it with
python -m vaani_observer.drain.
By design — it never raises, so an observability misconfiguration cannot take down your agent. The cost is silence. Assert at startup:
recorder = VaaniLiveKitRecorder.from_env(agent_id="support-bot")
if not recorder.enabled:
logger.warning("Vaanieval recording is disabled")The call appears but the trace is empty
An operation is written when it ends, not when it starts. An operation that
is never end()ed is deliberately absent from the package rather than persisted
as a span with no outcome, which would skew every duration percentile.
- Ensure every
start_operation()has a matchingend(). session.end()closes anything still open — a hard process kill does not.- With the LiveKit integration, pass
job_ctx=ctxtoobserve_agent_session()so the recording is finalized in the job's shutdown hook. Do not callfinish()in afinallyaroundsession.start(): on livekit-agents 1.xstart()returns when the session starts, so that truncates the call.
Two conditions must both hold, and both fail silently:
- An ambient session must be installed. Wrap the call in
session.context()/session.run(),with_turn/withTurn, orsession.bind(). - The URL must match an endpoint rule. Verify with:
print(vaani.classify_url("https://api.openai.com/v1/chat/completions"))console.log(vaani.classifyUrl('https://api.openai.com/v1/chat/completions'));None / null means no rule matched. Remember match: "path" is a prefix
match on the path, and the host must be identical.
Ambient context carries the session but not necessarily the turn. Scope
explicitly with session.with_turn(turn.id) / session.withTurn(turn.id, fn), or
attach late with op.set_turn(turn.id) / op.setTurn(turn.id).
With LiveKit, this is what VaaniAudioTapMixin does via llm_node — if you
subclassed Agent without the mixin, or forgot agent.vaani = recorder, LLM
spans will land without a turn.
Two rules match the same URL at the same scheme precedence. This is intentional — Vaanieval fails loudly rather than attributing your LLM latency to your TTS budget.
Fix by narrowing one rule to match: "exact", or by making the paths disjoint.
Note that a rule written for the exact scheme wins over the transport-neutral
form, so an https:// and a wss:// rule on the same host do not conflict.
Audio problems
capture.audiomay beFalse. The audio methods returnFalserather than raising when capture is off or the session has ended.- Python: a non-bytes chunk raises
TypeError; a wrong encoding, sample rate or channel count raisesValueError. - Node: all of the above raise
TypeError. - The format cannot change within a session, per track.
- LiveKit: audio comes from
VaaniAudioTapMixin. It must be listed beforeAgentin the base classes andagent.vaanimust be set.
That would mean agent chunks were placed at arrival time. The SDK avoids it with a playout clock: the agent track advances by the PCM duration of each chunk, so gaps between streamed TTS chunks are preserved.
If pauses are still missing, check that you are passing a correct
sample_rate_hz — an inflated rate makes each chunk appear shorter than it is,
compressing the timeline.
It has no container and no header — it is raw interleaved stereo PCM.
ffplay -f s16le -ar 16000 -ch_layout stereo call.audio
ffmpeg -f s16le -ar 16000 -ac 2 -i call.audio call.wavThe sample rate is in manifest.json under audio.call.sample_rate_hz.
The dashboard's WAV preview honours HTTP Range, which Safari requires before it
will play any media response. If you are proxying the dashboard, make sure your
proxy forwards Range and does not buffer the response.
Upload problems
Both must be set on the observer. If you only want local spooling, do not call
it — session.end() alone writes a complete package.
The idempotency-key header must be exactly the session_id. Any other
value is rejected.
The object exceeds 128 MiB. Raw stereo PCM at 16 kHz is roughly 3.8 MB per minute, so this is a very long call. There is no chunked upload path — reduce the sample rate, disable audio capture for that agent, or split the call.
gzip does not help here: the cap is enforced on decompressed bytes.
The upload is bounded twice: upload.timeout_s (default 30.0) per request, and
the timeout= argument to upload_package() for the whole upload. Retries cover
5xx, 408, 425, 429 and dropped connections.
A failed upload normally leaves the package on the spool and logs its directory. This is not a durability guarantee: an ephemeral host or hard kill can lose it. Retry from persistent storage:
python -m vaani_observer.drain --verboseOn LiveKit, an upload inside the job shutdown hook competes with
shutdown_process_timeout, which defaults to 10 seconds. Raise it, or leave
the upload to the drainer. See
LiveKit → Uploading takes time.
415 means a Content-Encoding the server does not implement was sent. gzip,
identity and an absent header are all accepted; anything else is a 415.
400 means the body claimed to be gzip and was malformed — most often a proxy
that decompressed the body but left the header in place.
Set VAANI_UPLOAD_COMPRESS=0 (or upload={"compress": False}) if something
between the SDK and the dashboard rewrites bodies.
The SDK only compresses when the ingest advertised gzip in
accepted_encodings on the 201, so against a conforming server neither of
these should be reachable. Seeing them points at something in the middle.
The server verifies both byte_size and sha256 for every declared object.
A mismatch means the upload was truncated or the digest was computed over
different bytes. Recompute from the exact file you uploaded:
wc -c < events.jsonl
shasum -a 256 events.jsonlThis check is why a truncated upload is a hard failure rather than a call with quietly missing turns.
The objects verified, but no operations could be read from events.jsonl.
partial is what surfaces as unverifiable in the dashboard.
Check the capture warning on the call first — it distinguishes the three causes:
| Warning says | Cause |
|---|---|
| the recorder measured 0ms of agent audio | Your agent never spoke. Nothing is wrong with the recording; debug the agent. |
| the agent spoke but nothing was classified | events.jsonl did not reach the observer, or every operation was left unended. |
| neither (an older SDK) | Usually a crash before session.end(). |
The measurement comes from the recorder's own tap in tts_node, so it does not
depend on any provider emitting metrics.
The uploader does not rewrite or repair existing manifests, and packages remain on the spool. However, an older manifest declaring less time than its rendered PCM is not rejected solely for that mismatch: the current hosted development backend accounts for the greater of declared duration and rounded-up PCM duration, preserving the original manifest and hash.
The unreleased checkout finalizers now round up the maximum of elapsed time, recorded event ends and rendered PCM duration for new sessions. This covers queued TTS tails and explicit future operation ends, without changing PCM bytes. An immediate empty session can still have zero duration.
The hosted effective-duration limit is 600,000 ms (10 minutes). Exceeding it or
another create-contract bound can return 422. Separately, import validates
operation timestamps against the effective timeline and audio-chunk ends against
actual PCM duration; incoherent events can fail the queued import.
Check the validation or import reason and retain the original package for diagnosis. Blind retries do not repair invalid evidence. These are hosted beta checks, not a claim that the local API uses the same validation. Use the pinned hosted SDK source previews in the quickstarts; older revisions may not implement this protocol.
Metrics are missing
The engine requires caller stops speaking and provider marks speech final to
have been observed separately. Many frameworks stamp both from the same
underlying event, making them byte-identical; subtracting them would manufacture a
zero, so the metric is reported as unavailable instead.
This is correct behaviour, not a bug. If you need the metric, the recognizer integration has to emit the two moments independently.
A turn missing any required milestone is excluded from percentiles rather than estimated. Common causes:
- A provider wrapper that never calls
op.event(...). - A code path that bypasses your endpoint rules, so no span is created.
metrics_collectedunavailable in your LiveKit version, removing TTFT/TTFB.
Expand a turn in the trace and see which milestone is absent.
A session-long STT websocket is being scored as a turn. Open it with
scope: "connection" — observe_websocket() / observeWebSocket() does this for
you. Per-turn STT timing must come from explicit operations.
Estimated WER requires a usable challenger comparison, not just timing capture.
Local evaluations default off; explicitly set VAANI_EVALUATIONS_ENABLED=1,
then read GET /v1/evaluation-policy. The UI's Run comparison requires
confirmation; API callers submit the returned consent_version with
model: "elevenlabs_scribe_v2".
It requires ELEVENLABS_API_KEY, transcript capture (stt_content) and recorded
caller audio. Full caller audio goes to ElevenLabs; transcript judging uses
OpenAI when configured, potentially with a compatible-model fallback. This can
incur provider charges and is not a ground-truth accuracy measurement.
503: evaluation is disabled. Provider keys alone do not enable it.409: the consent version is missing or stale. Fetch the current policy and review its disclosure again before confirming.422: the model key is unsupported; currently useelevenlabs_scribe_v2.
Check whether the operation ended with a cancellation. AbortError,
CancelledError and CancelledException are excluded from failure counts — a TTS
span aborted by barge-in is correct. If your provider raises a differently named
cancellation, end the span with status="cancelled" explicitly rather than
letting it record as an error.
Operational problems
Spool directories are not removed by upload_package(). Either add a sweep of
.vaani-spool for directories containing a manifest.json older than your
retention window, or run the drainer, which uploads what is pending and removes
what the dashboard has verified:
python -m vaani_observer.drain --watch 300A 5-minute call is roughly 28 MB of raw PCM, so an agent host recording all day fills a disk quickly.
The case that catches people out is the package the recorder itself already
uploaded from finish(). It carries an uploaded.json receipt, so the drainer
correctly refuses to re-ship it — and for a long time nothing removed it either,
so a healthy, fully-uploaded worker still filled its disk. Those packages are
now removed once their receipt is older than --delivered-ttl hours
(default 24).
# Keep delivered packages for an hour instead of a day
python -m vaani_observer.drain --watch 300 --delivered-ttl 1
# Never remove anything
python -m vaani_observer.drain --watch 300 --keepUse --keep if you want every package left behind with its receipt.
Age is read from the uploaded_at field inside uploaded.json, not from the
file's modification time, because anything that copies a spool — cp -R,
rsync without -a, a container image build, a backup restore — resets mtime
and would make every package look brand new forever.
There is no retention policy. Delete session directories under
$VAANI_DATA_DIR/objects and their rows in vaani.db. There is no deletion
API. Plan cleanup of related evaluation artifacts, SDK spools and backups too;
removing local data does not establish deletion at external providers.
Challenger jobs run on a two-worker in-process thread pool, so they compete with request serving. Queue fewer at a time, or run evaluation against a separate instance pointed at a copy of the data.
Jobs that were queued or in_progress when the process stopped are marked
failed on startup and are not resumed — a half-finished run is never presented
as complete. Re-queue them.
Shared-directory multi-instance operation is unsupported. Processes contend on SQLite and shared object files; this is not a horizontal scaling strategy. Run one localhost instance.
Still stuck
Dashboard repository
GitHub source repository. Open an issue with a synthetic reproduction and reviewed, minimal diagnostic details.
Email shubham@vaanieval.com
Ask for setup help or discuss a safe reproduction. Do not email raw customer audio, transcripts or credentials.
Book a 20-minute call
Walk through it with us against your own traffic.