Vaanieval
Reference

Session package

The shared on-disk session format — manifest, events and enabled audio — with capture-coverage and measurement caveats.

session.end() produces a directory. That directory is the contract between the capture SDKs and the dashboard, and both SDKs emit a byte-compatible format.

You can inspect it, diff it, archive it, or replay it into a different backend entirely.

manifest.json
events.jsonl
call.audio

Write invariants

Use the manifest as the finalization marker, then check coverage:

1. manifest.json is written last, using a staged write and rename. A directory without it is still being written or was abandoned.

2. Finalized does not mean fully captured. Inspect capture_status, configured capture settings and the recorded evidence. Upload digests verify file transfer, not that every live turn or audio frame was recorded.

manifest.json

{
  "schema_version": "1.0",
  "sdk": { "name": "@vaanieal/observer", "version": "0.1.0" },
  "session_id": "9f2c1e4a-3b7d-4c88-9a21-0e5f7b2d1c34",
  "agent_id": "support-bot",
  "metadata": { "env": "prod", "version": "2026.08.1" },
  "started_at": "2026-08-14T09:12:03.418Z",
  "duration_ms": 184320,
  "outcome": "completed",
  "capture_status": {
    "events_complete": true,
    "audio_complete": true,
    "http_instrumentation": "active",
    "websocket_instrumentation": "active",
    "dropped_event_count": 0,
    "dropped_audio_chunk_count": 0
  },
  "audio": {
    "call": {
      "file": "call.audio",
      "encoding": "pcm_s16le",
      "sample_rate_hz": 16000,
      "channels": 2,
      "channel_layout": { "left": "agent", "right": "caller" }
    }
  }
}

Prop

Type

duration_ms and recorded tails

Unreleased development-checkout behavior: new Python and Node.js session finalizers round up to a whole millisecond after taking the maximum of:

  • Elapsed session time.
  • Recorded event timeline ends, including explicitly future-dated operation ends.
  • Actual rendered PCM duration, including queued or scheduled TTS tails.

The manifest therefore covers recorded output that extends beyond the instant end() was called. It is not necessarily the telephone connection's wall-clock duration. PCM bytes are unchanged by this duration correction, and it adds no new audio metadata field. An immediate, empty synthetic session without audio can still have duration_ms: 0.

A larger declared duration can increase duration-based hosted minute reservations; this is not a provider invoice or a quote for evaluation costs. These finalizer changes are working-tree-only, not available in the currently pushed public Git revisions.

Existing finalized packages are not repaired during upload. Their manifest is preserved. An older manifest that understates its PCM duration is not automatically rejected for that mismatch by the current hosted development backend. For reservation and usage accounting, that backend separately derives:

pcm_ms = ceil(call.audio byte_size * 1000 / (sample_rate_hz * channels * 2))
effective_duration_ms = max(original manifest.duration_ms, pcm_ms)

Without audio, the effective duration is the declared duration. The hosted development limit is 600,000 ms (10 minutes); exceeding the supported duration or another create-contract bound can produce 422. The original manifest and its hash remain unchanged.

During import, operation timestamps must fit the effective timeline, while audio_chunk ends must fit the actual PCM recording. Invalid event evidence can still fail import; extending accounting to cover PCM does not repair arbitrary event timestamps or prove complete capture. The uploader retains packages regardless of success or failure. These are development-checkout semantics, not a published hosted compatibility guarantee.

capture_status

This is how degraded capture becomes visible when strict is off.

FieldMeaning
events_completeNo event write failed
audio_completeNo audio chunk write failed
http_instrumentation"active" when HTTP patching is installed, else "disabled"
websocket_instrumentation"active" when websocket observation is enabled, else "disabled"
dropped_event_countEvents lost to write failures
dropped_audio_chunk_countAudio chunks lost to write failures
coverage_completefalse when any reply is not fully accounted for by an operation
coverage_gapsOne entry per proven gap, naming the stage, the reason and the turns
measuredWhat the recorder measured for itself, independently of any plugin

A gap entry always carries stage, reason and turn_ids; the remaining fields depend on what the gap can actually support. Two reasons are easy to confuse and are deliberately reported apart:

  • "agent audio was rendered that no tts operation accounts for" — the caller heard something the spans do not describe. It carries unattributed_agent_audio_ms, and latency and cost for those turns are understated by a measurable amount.
  • "a reply's text was captured but it was never rendered as audio" — the reply was composed but nothing was played, usually because the call ended while the agent was still speaking. There is no unaccounted audio to report, so no audio total is published with it.

capture_status.measured

The LiveKit integration tees every frame on its way through tts_node, so it holds proof of what the agent actually said that does not depend on any plugin reporting anything. That is what makes "zero TTS spans" answerable: a silent agent and a failed capture look identical without it.

FieldMeaning
agent_audio_msAgent speech measured from the frames themselves
agent_audio_tappedfalse when no tap was installed, so nothing here was verified
derived_tts_op_countSpans rebuilt by the SDK because the TTS plugin reported no metrics
derived_tts_agent_audio_msSpeech covered by those rebuilt spans
derived_tts_share_pctShare of the agent's speech described by rebuilt spans
reconstructed_op_countSpans rebuilt from the captured audio alone, with no report from any stage
tail_written_off_msMeasured speech forgiven as turn-boundary jitter
tail_write_off_cap_msThe absolute ceiling on that write-off for any call
tail_written_off_turn_idsWhich turns those forgiven milliseconds sat on
unattributed_agent_audio_msMeasured speech on no span and not written off
unattributed_tolerance_msHow much of that is tolerated before the call is called incomplete
stream_ownershipproved when every reply's audio was matched to it by identity, inferred when by timing

A rebuilt span is not a lost one. Every millisecond is still attributed, each span carries request.derived_from and response.estimated_fields, and derived_tts_share_pct says how much of the page is an estimate — on Deepgram aura-2, roughly three replies in four emit no metric, so this is routinely high on a perfectly healthy call. coverage_complete answers a different question: whether any speech is missing from the record entirely.

unattributed_agent_audio_ms is published on every call, including when it is zero and when it sits under unattributed_tolerance_ms. A threshold that is applied but never shown is a second write-off stacked on the first, so the number is always there to be audited even when it does not move the status.

tail_written_off_ms only ever covers audio that arrived without a stream token — frames whose reply was chosen by timing at a turn boundary. Audio rendered through tts_node names the reply that produced it, so a residual on one of those streams is a lifecycle defect rather than jitter and is reported as unattributed instead of forgiven. Two further limits: only the part that arrived after the reply's span was published can be forgiven, since anything earlier is already inside its played_ms; and a reply whose span was ended by a session error is never eligible, because such a span publishes no accounting to measure a residual against.

When a reply has no provider character count

LiveKit tags each llm and tts metric with the id of the speech that produced it, taken from its own speech-handle context. When that tag is absent the framework itself could not say which reply the measurement belongs to, and neither can we: both stages are additive — one reply emits one metric per synthesis segment or tool-call round — so "the previous reply already reported" never proves it has finished reporting.

Rather than guess, the SDK refuses the metric once a second reply exists, adds a coverage_gaps entry naming the stage, and rebuilds the reply's span from the captured audio and the transcript with response.estimated set. The reply, its words and its measured duration all survive; what is lost is the provider's own character count, and cost is billed off that. We would rather publish a disclosed gap than bill one reply for another reply's characters, which is a wrong invoice that looks exactly like a right one.

stream_ownership records how that matching was done. Normally the SDK reads LiveKit's speech-handle context, so each tts_node call names the reply it is rendering and attribution is an identity lookup. On a build that does not expose that context the SDK falls back to pinning the stream when tts_node is invoked, which is sound but weaker: per-turn talk time, latency and cost can move between adjacent replies, while call totals stay correct. The field says which of the two a per-turn number rests on, and the dashboard warns when it is inferred.

Per-span fields worth knowing when a number looks surprising:

FieldMeaning
response.played_msDuration observed in captured agent output; not independent proof of caller delivery
response.audio_msWhat the provider says it synthesized
response.provider_audio_ms_undercount_msSet when the caller heard more than the provider billed
response.ended_at_source"played_audio" when the span was extended to cover playout
response.text_source"tts_node" when the words are what we asked for rather than what was confirmed
response.segment_countProvider metrics summed into this span

On an STT span:

FieldMeaning
response.final_segmentsFinals the provider issued for this one committed message
response.provider_metered_audio_msThe recogniser usage that arrived while this turn was open. This is placement, not invoice attribution — for a connection-metering provider it is an arbitrary slice of a session-wide meter. See below
response.input_tokens / response.output_tokensRecogniser token usage, when the provider reports it (OpenAI does; Deepgram does not)
response.metered_after_finalSome of the meter arrived after this turn's transcript was final. It stays on this turn rather than moving to whoever spoke next, which is a placement rule, not a claim that this turn incurred it
response.metered_after_final_msHow much of the published meter arrived after the caller stopped talking
response.metered_arrivalWhen the meter arrived, relative to the final transcript: before_final, after_final or straddles_final. An observation about timing only
response.metering_scopeAlways unknown. See below — the SDK will not guess what the meter measures
response.metering_scope_noteThe caveat, carried on the payload itself so it travels with the number into whatever reads it
response.continues_turnPresent only when we knowingly recorded two turns where LiveKit committed one message; names the turn this one continues, with split_reason saying why they could not be merged. The call rollup counts such a pair once and reports split_turn_count, so turn_count keeps agreeing with LiveKit's own history
response.reply_includes_fillerThe reply's audio includes a say() filler spoken before the generated answer; filler_audio_ms gives the filler's share so the answer's own duration stays recoverable
response.filler_audio_ms_unknownThe filler ran but its meter had not reported when the answer began. LiveKit emits speech_created for a say() before scheduling its TTS, so this is an ordinary race — the filler is declared, only its duration is missing
response.reply_attributionPresent, and always "inferred", when this reply was kept in a turn of its own on evidence that could not settle the question. See below
response.reply_attribution_reasonPlain-English statement of what could not be distinguished, so you can judge the call rather than take our word for it
response.reply_skippedPresent when no reply is coming: "stop_response" (the agent ignored the utterance) or "callback_error" (its turn callback raised and LiveKit swallowed it)

What provider_metered_audio_ms measures

It is the number the recogniser will bill, summed over every increment it reported while this turn was open. It is not the caller's speech duration, and how far from it you are depends on which recogniser you run:

  • OpenAI emits the item's final transcript and then that item's usage, computed from the item's own start and end. That is a per-utterance meter — and it arrives after the final.
  • Deepgram pushes every frame streamed to the socket into a five-second collector and emits each tick. That is a connection meter, silence included, and a tick routinely lands before the final while the caller is still speaking. It also lands on whichever turn happens to be open, so it can exceed the span that carries it.

Those two cases are timed in opposite ways to what you would guess, so an earlier version of this SDK — which classified the meter as utterance or connection from its arrival time — was confidently wrong for both. It no longer guesses. metered_arrival reports the timing, which is observable; metering_scope reports unknown, because the metering rule belongs to the provider and is not in the payload.

Do not derive per-turn cost from this field. For a connection-metering provider such as Deepgram, the increments carry no turn identity at all: the SDK can tell you what a call was metered, and which turn happened to be open when each increment landed, but not which turn incurred it. Summing the field across a call is sound; dividing it by turn is false precision. For an utterance-metering provider the per-turn reading is meaningful — but the payload cannot tell you which kind you have, which is exactly what metering_scope: "unknown" is saying.

For how long the caller actually spoke, use presentation_window or the turn's user_speech_ms — those are measured from the audio and carry no such caveat.

When reply_attribution is inferred

If your agent speaks a filler from inside on_user_turn_completed — "let me check that for you" before a slow lookup — and a partial transcript arrives while that filler is playing, the reply that follows may be the answer to the caller whose filler is playing, or the answer to whoever just started speaking over it. The user_input_transcribed(is_final=False) event is the same either way, so the events alone cannot separate them.

The speech handle and the speech_created event settle three of the ways LiveKit creates a reply. The rest emit an event identical in every field to the automatic answer, so they are settled by what the recorder observed instead — specifically, by which frame asked for the reply:

How the reply was createdWhat it answersRecorded as
The automatic answer to a completed user turn — scheduled at once, user_initiatedThe turn that just finishedMerged into that turn, no flag
A preemptive generation, started from a predicted end of speech — stays unscheduled until LiveKit validates the predictionThe caller now speakingIts own turn, no flag
A realtime model generating server-side while it transcribes (agent_activity.py:1991-2008) — reports source generate_reply like the automatic answer, but is scheduled without being user_initiated, never passes through the public method, and leaves no tool result outstandingThe speech being transcribed now, whose final transcript may arrive after the replyIts own turn, joined by that transcript, no flag
The automatic reply after a tool result — also not user_initiated, but the current turn has a function_tools_executed result still unansweredThe turn that called the toolMerged into that turn, reply_attribution: "inferred"
That same reply arriving after LiveKit stopped waiting for it — its five-second wait (agent_activity.py:4278) bounds the wait, not the provider's generationEither the turn that called the tool or the next caller; the event says neitherIts own turn, reply_attribution: "inferred"
AgentSession.generate_reply() called by your own code — recognised because the call is still on the stack when the event fires, and the frame that made it is outside the livekit packageUnknowableIts own turn, reply_attribution: "inferred"
AgentSession.run(user_input=…) — LiveKit's own file, but an entry point only your code reachesThe user_input you supplied; no caller spokeIts own turn, reply_attribution: "inferred"
AgentSession.generate_reply() called by LiveKit itself — committing a realtime turn, the IVR activity, or a beta/workflows promptUsually something already under way, but the workflow and IVR helpers prompt proactivelyIts own turn, joined by the following transcript, reply_attribution: "inferred"
That same internal call made while a tool result on the turn is still owed a reply — LiveKit's asynchronous tool executor does exactly this (voice/tool_executor.py:589-603)The turn that ran the tool, most likelyMerged into that turn, reply_attribution: "inferred"
LiveKit reissuing a reply for a run that already exists — RunResult._maybe_retry_output() (run_result.py:292) after a structured output fails to validate, or realtime_fallback_adapter.py:394 when a realtime session drops to a text modelWhoever started that run; the event carries no run identity to say whoIts own turn, not joined by the following transcript, reply_attribution: "inferred"
The public method was replaced — a subclass or wrapper stands in for generate_reply, so neither it nor AgentActivity._generate_reply() is on the stack with LiveKit above itUnknowableIts own turn, reply_attribution: "inferred"
The stack could not be read — introspection blocked, the public method's code object unreachable, or livekit.agents imports but its modules have moved on this buildUnknowableIts own turn, reply_attribution: "inferred"
A decorated or traced generate_reply — functools.wraps copies the name but not the code object, so the stand-in frame is literally called wrapper and the override reaches a different internal pathUnknowableIts own turn, reply_attribution: "inferred"
A session that is not LiveKit's — a compatible or test double passed to attach() on a machine where livekit.agents is not importable at allUnknowableIts own turn, reply_attribution: "inferred"

Two rows stay judgements: the post-tool reply, and your own generate_reply(). Yours is scheduled and user_initiated exactly like the automatic answer — passing input_modality="audio" makes it identical — and your application may call it to answer the caller or to say something unrelated, so the reply is kept in a turn of its own.

The distinction between your call and LiveKit's is read from the call stack rather than from a replaced method, which matters in two everyday cases: saving the bound method (reply = session.generate_reply) before recording starts does not hide the call, and LiveKit's own internal uses of the method — more than a dozen on 1.7.0 — are not reported as yours. A subclass or wrapper that replaces the method entirely is recognised too, but not by looking for it. LiveKit's automatic answer is emitted from inside AgentActivity._generate_reply() (agent_activity.py:1559), synchronously, so that frame is on the stack every single time the automatic answer is created. Its absence therefore rules the automatic answer out, whatever the reply turns out to be — which is what makes a decorated override safe: functools.wraps copies the name, so a wrapper frame is still literally called wrapper, and no amount of name matching would find it. A frame calling itself generate_reply on one of the recorded sessions is used as a weaker second reading, for builds where that anchor cannot be resolved at all — and attach() may be called more than once, so every session it was given counts, not only the most recent one.

Where livekit.agents cannot be imported at all, none of this evidence exists and there is also no LiveKit to have generated anything — so a reply that cannot be traced to the recorded session's own call is reported as unreadable rather than assumed to be automatic. The cost is a caveat on a compatible framework that generates replies by itself; the alternative was attributing those replies to whoever spoke last and saying nothing.

LiveKit's own calls are still flagged, because not all of them answer speech already under way: the beta/workflows and IVR helpers ask the caller a question before the caller has said anything, and the transcript that follows is the answer to it rather than the question it answered. The events are identical either way. The post-tool reply is placed on the turn that ran the tool, which is where it almost always belongs, but a caller talking over the tool call could have prompted it instead. Either way the flag is set on every span of that turn, including the LLM span, so an interrupted reply that reported tokens still carries the caveat.

The caveat is also counted, not only recorded. The dashboard's fleet view reports coverage.inferred_reply_turns for the range you are looking at, and says so on the page when it is non-zero. Without that the flag reached the archive and stopped there: a range's latency, token and cost figures could move because a reply had been placed by reading the call rather than by the events, and nothing on the page would have mentioned it. The turn is not subtracted from the turn count — the exchange happened and its measurements are real; what is in doubt is which turn they belong to.

If you are running a build of livekit-agents that does not expose these signals, no reply is merged and the flag is set — the recorder never guesses silently because a field was missing. The same applies to the call stack: a failure to read it is not treated as evidence that the reply was automatic.

Treat a turn carrying this flag as unreliable for response latency specifically. Everything else on it — audio, durations, token counts — is measured, not inferred.

reply_skipped separates a turn the agent chose not to answer from one whose reply failed to sound. Both look identical in the numbers — a transcript, no TTS span — and they call for opposite responses, so the absence of a reply is never by itself reported as a declined turn. It is set only when the agent said so, and requires agent= on observe_agent_session.

A caller's speech window ends at speech_ended, which the voice detector stamps at the first pause. When one message arrives as several finals, the later ones are evidence the caller kept talking, so the window ends at the last final less the recogniser's own transcription_delay_ms, and presentation_window.source reads final_transcripts to say so.

A session with events_complete: false or a non-zero drop count is classified unverifiable in the dashboard, not healthy and not failed. "We do not know what the caller heard" is a distinct claim, and it is ranked above merely slow calls precisely because it hides problems.

events.jsonl

One JSON object per line, in append order. Operation spans are written when they end and are identified by their type; audio and websocket lifecycle events carry a kind field instead (audio_chunk, websocket, and — in Python — capture_error).

operation

Written when the operation ends.

{
  "event_id": "b41c…",
  "session_id": "9f2c…",
  "turn_id": "turn-2",
  "scope": "turn",
  "type": "llm",
  "endpoint_id": "llm",
  "provider": "openai",
  "model": "gpt-4o-mini",
  "transport": "http",
  "started_at_ms": 4821,
  "ended_at_ms": 6104,
  "duration_ms": 1283,
  "status": "ok",
  "request": { "messages": "…" },
  "response": { "tokens": 128 },
  "error": null,
  "milestones": {
    "first_token": { "occurred_at_ms": 5290, "last_at_ms": 5290, "count": 1 }
  },
  "samples": {
    "partial_transcript": { "items": [{ "occurred_at_ms": 5010, "text": "how do i" }], "truncated": false }
  }
}
FieldNotes
typestt | llm | tts | tool
scopeturn (default) or connection
transportmanual, or set by instrumentation (http, websocket)
statusok by default; anything else marks the span failed
milestonesRepeated names merge: first occurred_at_ms, latest last_at_ms, count
samplesPer-name bucket, capped at limit (default 100), with a truncated flag

audio_chunk

{ "kind": "audio_chunk", "track": "caller", "occurred_at_ms": 1240, "byte_length": 3200, "duration_ms": 100 }

track is caller or agent. duration_ms is present when the PCM duration is computable. The agent track's occurred_at_ms comes from the playout clock, not arrival time — see below.

websocket

{ "kind": "websocket", "session_id": "9f2c…", "occurred_at_ms": 980, "event": "closed", "url": "wss://…", "code": 1000 }

Payload bounding

Any captured payload larger than payload_max_bytes (16 KiB default) is replaced in place:

{ "_truncated": true, "_original_bytes": 184320, "_preview": "{\"messages\":[{\"role\"…" }

A value that cannot be serialised becomes { "_capture_error": "…" }. The event is still written — capture never fails an operation.

call.audio

Raw interleaved stereo PCM. No container, no header.

PropertyValue
Encodingpcm_s16le
Channels2
Sample rateThe maximum of the two input track rates
Left channelAgent
Right channelCaller
Frame size4 bytes

Bytes are frames × 4, where frames is the larger of the session duration in samples and the longest rendered track. A track with no audio is silence.

The playout clock

TTS audio usually arrives in a burst even though it will be played in real time. If chunks were placed at arrival time, a 12-second reply would appear as a half-second blob and every pause in the call would vanish.

Instead, the agent track's clock advances by the PCM duration of each chunk: a chunk is placed at max(arrival_time, end_of_previous_chunk). That preserves the real silences, which is what makes the waveform's gap annotations meaningful and what makes latency audible when you scrub to a timestamp.

Playing it

ffplay -f s16le -ar 16000 -ch_layout stereo call.audio
ffmpeg -f s16le -ar 16000 -ac 2 -i call.audio call.wav

The dashboard does the same thing on demand — it wraps the raw PCM in a WAV header per request rather than storing a second copy, and honours HTTP Range because Safari refuses to play media otherwise.

Audio dominates storage. Two raw 16-bit PCM tracks are roughly 64 KB per second of call at 16 kHz — about 3.8 MB per minute. Stored audio is raw PCM; transfer compression does not shrink stored objects. Upload methods do not remove spool copies, though the Python drainer has separate cleanup settings. The local dashboard has no automated retention policy. Plan disk capacity, backup and cleanup before increasing call volume.

Consuming a package yourself

The format is deliberately boring, so you do not have to use the dashboard:

import json
from pathlib import Path

package = Path("./.vaani-spool/9f2c1e4a-…")
manifest = json.loads((package / "manifest.json").read_text())

operations = [
    json.loads(line)
    for line in (package / "events.jsonl").read_text().splitlines()
    if line and json.loads(line).get("type") in {"stt", "llm", "tts", "tool"}
]

turns = {}
for op in operations:
    if op.get("scope") == "connection":
        continue          # a session-long socket is not an utterance
    turns.setdefault(op["turn_id"], []).append(op)

Two rules to honour if you build on this: skip scope: "connection" spans when computing per-turn metrics, and treat a missing milestone as unmeasurable rather than substituting an operation's start or end time. Both are the reason the dashboard's numbers hold up.

Next

On this page