Session package
The shared on-disk session format — manifest, events and enabled audio — with capture-coverage and measurement caveats.
session.end() produces a directory. That directory is the contract between the
capture SDKs and the dashboard, and both SDKs emit a byte-compatible format.
You can inspect it, diff it, archive it, or replay it into a different backend entirely.
Write invariants
Use the manifest as the finalization marker, then check coverage:
1. manifest.json is written last, using a staged write and rename.
A directory without it is still being written or was abandoned.
2. Finalized does not mean fully captured. Inspect capture_status,
configured capture settings and the recorded evidence. Upload digests verify
file transfer, not that every live turn or audio frame was recorded.
manifest.json
{
"schema_version": "1.0",
"sdk": { "name": "@vaanieal/observer", "version": "0.1.0" },
"session_id": "9f2c1e4a-3b7d-4c88-9a21-0e5f7b2d1c34",
"agent_id": "support-bot",
"metadata": { "env": "prod", "version": "2026.08.1" },
"started_at": "2026-08-14T09:12:03.418Z",
"duration_ms": 184320,
"outcome": "completed",
"capture_status": {
"events_complete": true,
"audio_complete": true,
"http_instrumentation": "active",
"websocket_instrumentation": "active",
"dropped_event_count": 0,
"dropped_audio_chunk_count": 0
},
"audio": {
"call": {
"file": "call.audio",
"encoding": "pcm_s16le",
"sample_rate_hz": 16000,
"channels": 2,
"channel_layout": { "left": "agent", "right": "caller" }
}
}
}Prop
Type
duration_ms and recorded tails
Unreleased development-checkout behavior: new Python and Node.js session finalizers round up to a whole millisecond after taking the maximum of:
- Elapsed session time.
- Recorded event timeline ends, including explicitly future-dated operation ends.
- Actual rendered PCM duration, including queued or scheduled TTS tails.
The manifest therefore covers recorded output that extends beyond the instant
end() was called. It is not necessarily the telephone connection's wall-clock
duration. PCM bytes are unchanged by this duration correction, and it adds no new
audio metadata field. An immediate, empty synthetic session without audio can
still have duration_ms: 0.
A larger declared duration can increase duration-based hosted minute reservations; this is not a provider invoice or a quote for evaluation costs. These finalizer changes are working-tree-only, not available in the currently pushed public Git revisions.
Existing finalized packages are not repaired during upload. Their manifest is preserved. An older manifest that understates its PCM duration is not automatically rejected for that mismatch by the current hosted development backend. For reservation and usage accounting, that backend separately derives:
pcm_ms = ceil(call.audio byte_size * 1000 / (sample_rate_hz * channels * 2))
effective_duration_ms = max(original manifest.duration_ms, pcm_ms)Without audio, the effective duration is the declared duration. The hosted
development limit is 600,000 ms (10 minutes); exceeding the supported
duration or another create-contract bound can produce 422. The original
manifest and its hash remain unchanged.
During import, operation timestamps must fit the effective timeline, while
audio_chunk ends must fit the actual PCM recording. Invalid event evidence can
still fail import; extending accounting to cover PCM does not repair arbitrary
event timestamps or prove complete capture. The uploader retains packages
regardless of success or failure. These are development-checkout semantics, not
a published hosted compatibility guarantee.
capture_status
This is how degraded capture becomes visible when strict is off.
| Field | Meaning |
|---|---|
events_complete | No event write failed |
audio_complete | No audio chunk write failed |
http_instrumentation | "active" when HTTP patching is installed, else "disabled" |
websocket_instrumentation | "active" when websocket observation is enabled, else "disabled" |
dropped_event_count | Events lost to write failures |
dropped_audio_chunk_count | Audio chunks lost to write failures |
coverage_complete | false when any reply is not fully accounted for by an operation |
coverage_gaps | One entry per proven gap, naming the stage, the reason and the turns |
measured | What the recorder measured for itself, independently of any plugin |
A gap entry always carries stage, reason and turn_ids; the remaining
fields depend on what the gap can actually support. Two reasons are easy to
confuse and are deliberately reported apart:
- "agent audio was rendered that no tts operation accounts for" — the caller
heard something the spans do not describe. It carries
unattributed_agent_audio_ms, and latency and cost for those turns are understated by a measurable amount. - "a reply's text was captured but it was never rendered as audio" — the reply was composed but nothing was played, usually because the call ended while the agent was still speaking. There is no unaccounted audio to report, so no audio total is published with it.
capture_status.measured
The LiveKit integration tees every frame on its way through tts_node, so it
holds proof of what the agent actually said that does not depend on any plugin
reporting anything. That is what makes "zero TTS spans" answerable: a silent
agent and a failed capture look identical without it.
| Field | Meaning |
|---|---|
agent_audio_ms | Agent speech measured from the frames themselves |
agent_audio_tapped | false when no tap was installed, so nothing here was verified |
derived_tts_op_count | Spans rebuilt by the SDK because the TTS plugin reported no metrics |
derived_tts_agent_audio_ms | Speech covered by those rebuilt spans |
derived_tts_share_pct | Share of the agent's speech described by rebuilt spans |
reconstructed_op_count | Spans rebuilt from the captured audio alone, with no report from any stage |
tail_written_off_ms | Measured speech forgiven as turn-boundary jitter |
tail_write_off_cap_ms | The absolute ceiling on that write-off for any call |
tail_written_off_turn_ids | Which turns those forgiven milliseconds sat on |
unattributed_agent_audio_ms | Measured speech on no span and not written off |
unattributed_tolerance_ms | How much of that is tolerated before the call is called incomplete |
stream_ownership | proved when every reply's audio was matched to it by identity, inferred when by timing |
A rebuilt span is not a lost one. Every millisecond is still attributed, each
span carries request.derived_from and response.estimated_fields, and
derived_tts_share_pct says how much of the page is an estimate — on
Deepgram aura-2, roughly three replies in four emit no metric, so this is
routinely high on a perfectly healthy call. coverage_complete answers a
different question: whether any speech is missing from the record entirely.
unattributed_agent_audio_ms is published on every call, including when it is
zero and when it sits under unattributed_tolerance_ms. A threshold that is
applied but never shown is a second write-off stacked on the first, so the
number is always there to be audited even when it does not move the status.
tail_written_off_ms only ever covers audio that arrived without a stream
token — frames whose reply was chosen by timing at a turn boundary. Audio
rendered through tts_node names the reply that produced it, so a residual on
one of those streams is a lifecycle defect rather than jitter and is reported
as unattributed instead of forgiven. Two further limits: only the part that
arrived after the reply's span was published can be forgiven, since anything
earlier is already inside its played_ms; and a reply whose span was ended by
a session error is never eligible, because such a span publishes no accounting
to measure a residual against.
When a reply has no provider character count
LiveKit tags each llm and tts metric with the id of the speech that
produced it, taken from its own speech-handle context. When that tag is absent
the framework itself could not say which reply the measurement belongs to, and
neither can we: both stages are additive — one reply emits one metric per
synthesis segment or tool-call round — so "the previous reply already reported"
never proves it has finished reporting.
Rather than guess, the SDK refuses the metric once a second reply exists, adds
a coverage_gaps entry naming the stage, and rebuilds the reply's span from
the captured audio and the transcript with response.estimated set. The reply,
its words and its measured duration all survive; what is lost is the provider's
own character count, and cost is billed off that. We would rather publish a
disclosed gap than bill one reply for another reply's characters, which is a
wrong invoice that looks exactly like a right one.
stream_ownership records how that matching was done. Normally the SDK reads
LiveKit's speech-handle context, so each tts_node call names the reply it is
rendering and attribution is an identity lookup. On a build that does not expose
that context the SDK falls back to pinning the stream when tts_node is
invoked, which is sound but weaker: per-turn talk time, latency and cost can
move between adjacent replies, while call totals stay correct. The field says
which of the two a per-turn number rests on, and the dashboard warns when it is
inferred.
Per-span fields worth knowing when a number looks surprising:
| Field | Meaning |
|---|---|
response.played_ms | Duration observed in captured agent output; not independent proof of caller delivery |
response.audio_ms | What the provider says it synthesized |
response.provider_audio_ms_undercount_ms | Set when the caller heard more than the provider billed |
response.ended_at_source | "played_audio" when the span was extended to cover playout |
response.text_source | "tts_node" when the words are what we asked for rather than what was confirmed |
response.segment_count | Provider metrics summed into this span |
On an STT span:
| Field | Meaning |
|---|---|
response.final_segments | Finals the provider issued for this one committed message |
response.provider_metered_audio_ms | The recogniser usage that arrived while this turn was open. This is placement, not invoice attribution — for a connection-metering provider it is an arbitrary slice of a session-wide meter. See below |
response.input_tokens / response.output_tokens | Recogniser token usage, when the provider reports it (OpenAI does; Deepgram does not) |
response.metered_after_final | Some of the meter arrived after this turn's transcript was final. It stays on this turn rather than moving to whoever spoke next, which is a placement rule, not a claim that this turn incurred it |
response.metered_after_final_ms | How much of the published meter arrived after the caller stopped talking |
response.metered_arrival | When the meter arrived, relative to the final transcript: before_final, after_final or straddles_final. An observation about timing only |
response.metering_scope | Always unknown. See below — the SDK will not guess what the meter measures |
response.metering_scope_note | The caveat, carried on the payload itself so it travels with the number into whatever reads it |
response.continues_turn | Present only when we knowingly recorded two turns where LiveKit committed one message; names the turn this one continues, with split_reason saying why they could not be merged. The call rollup counts such a pair once and reports split_turn_count, so turn_count keeps agreeing with LiveKit's own history |
response.reply_includes_filler | The reply's audio includes a say() filler spoken before the generated answer; filler_audio_ms gives the filler's share so the answer's own duration stays recoverable |
response.filler_audio_ms_unknown | The filler ran but its meter had not reported when the answer began. LiveKit emits speech_created for a say() before scheduling its TTS, so this is an ordinary race — the filler is declared, only its duration is missing |
response.reply_attribution | Present, and always "inferred", when this reply was kept in a turn of its own on evidence that could not settle the question. See below |
response.reply_attribution_reason | Plain-English statement of what could not be distinguished, so you can judge the call rather than take our word for it |
response.reply_skipped | Present when no reply is coming: "stop_response" (the agent ignored the utterance) or "callback_error" (its turn callback raised and LiveKit swallowed it) |
What provider_metered_audio_ms measures
It is the number the recogniser will bill, summed over every increment it reported while this turn was open. It is not the caller's speech duration, and how far from it you are depends on which recogniser you run:
- OpenAI emits the item's final transcript and then that item's usage, computed from the item's own start and end. That is a per-utterance meter — and it arrives after the final.
- Deepgram pushes every frame streamed to the socket into a five-second collector and emits each tick. That is a connection meter, silence included, and a tick routinely lands before the final while the caller is still speaking. It also lands on whichever turn happens to be open, so it can exceed the span that carries it.
Those two cases are timed in opposite ways to what you would guess, so an
earlier version of this SDK — which classified the meter as utterance or
connection from its arrival time — was confidently wrong for both. It no
longer guesses. metered_arrival reports the timing, which is observable;
metering_scope reports unknown, because the metering rule belongs to the
provider and is not in the payload.
Do not derive per-turn cost from this field. For a connection-metering
provider such as Deepgram, the increments carry no turn identity at all: the SDK
can tell you what a call was metered, and which turn happened to be open when
each increment landed, but not which turn incurred it. Summing the field across
a call is sound; dividing it by turn is false precision. For an
utterance-metering provider the per-turn reading is meaningful — but the payload
cannot tell you which kind you have, which is exactly what metering_scope: "unknown" is saying.
For how long the caller actually spoke, use presentation_window or the turn's
user_speech_ms — those are measured from the audio and carry no such caveat.
When reply_attribution is inferred
If your agent speaks a filler from inside on_user_turn_completed — "let me
check that for you" before a slow lookup — and a partial transcript arrives
while that filler is playing, the reply that follows may be the answer to the
caller whose filler is playing, or the answer to whoever just started speaking
over it. The user_input_transcribed(is_final=False) event is the same either
way, so the events alone cannot separate them.
The speech handle and the speech_created event settle three of the ways
LiveKit creates a reply. The rest emit an event identical in every field to the
automatic answer, so they are settled by what the recorder observed instead —
specifically, by which frame asked for the reply:
| How the reply was created | What it answers | Recorded as |
|---|---|---|
The automatic answer to a completed user turn — scheduled at once, user_initiated | The turn that just finished | Merged into that turn, no flag |
| A preemptive generation, started from a predicted end of speech — stays unscheduled until LiveKit validates the prediction | The caller now speaking | Its own turn, no flag |
A realtime model generating server-side while it transcribes (agent_activity.py:1991-2008) — reports source generate_reply like the automatic answer, but is scheduled without being user_initiated, never passes through the public method, and leaves no tool result outstanding | The speech being transcribed now, whose final transcript may arrive after the reply | Its own turn, joined by that transcript, no flag |
The automatic reply after a tool result — also not user_initiated, but the current turn has a function_tools_executed result still unanswered | The turn that called the tool | Merged into that turn, reply_attribution: "inferred" |
That same reply arriving after LiveKit stopped waiting for it — its five-second wait (agent_activity.py:4278) bounds the wait, not the provider's generation | Either the turn that called the tool or the next caller; the event says neither | Its own turn, reply_attribution: "inferred" |
AgentSession.generate_reply() called by your own code — recognised because the call is still on the stack when the event fires, and the frame that made it is outside the livekit package | Unknowable | Its own turn, reply_attribution: "inferred" |
AgentSession.run(user_input=…) — LiveKit's own file, but an entry point only your code reaches | The user_input you supplied; no caller spoke | Its own turn, reply_attribution: "inferred" |
AgentSession.generate_reply() called by LiveKit itself — committing a realtime turn, the IVR activity, or a beta/workflows prompt | Usually something already under way, but the workflow and IVR helpers prompt proactively | Its own turn, joined by the following transcript, reply_attribution: "inferred" |
That same internal call made while a tool result on the turn is still owed a reply — LiveKit's asynchronous tool executor does exactly this (voice/tool_executor.py:589-603) | The turn that ran the tool, most likely | Merged into that turn, reply_attribution: "inferred" |
LiveKit reissuing a reply for a run that already exists — RunResult._maybe_retry_output() (run_result.py:292) after a structured output fails to validate, or realtime_fallback_adapter.py:394 when a realtime session drops to a text model | Whoever started that run; the event carries no run identity to say who | Its own turn, not joined by the following transcript, reply_attribution: "inferred" |
The public method was replaced — a subclass or wrapper stands in for generate_reply, so neither it nor AgentActivity._generate_reply() is on the stack with LiveKit above it | Unknowable | Its own turn, reply_attribution: "inferred" |
The stack could not be read — introspection blocked, the public method's code object unreachable, or livekit.agents imports but its modules have moved on this build | Unknowable | Its own turn, reply_attribution: "inferred" |
A decorated or traced generate_reply — functools.wraps copies the name but not the code object, so the stand-in frame is literally called wrapper and the override reaches a different internal path | Unknowable | Its own turn, reply_attribution: "inferred" |
A session that is not LiveKit's — a compatible or test double passed to attach() on a machine where livekit.agents is not importable at all | Unknowable | Its own turn, reply_attribution: "inferred" |
Two rows stay judgements: the post-tool reply, and your own generate_reply().
Yours is scheduled and user_initiated exactly like the automatic answer —
passing input_modality="audio" makes it identical — and your application may
call it to answer the caller or to say something unrelated, so the reply is kept
in a turn of its own.
The distinction between your call and LiveKit's is read from the call stack
rather than from a replaced method, which matters in two everyday cases: saving
the bound method (reply = session.generate_reply) before recording starts does
not hide the call, and LiveKit's own internal uses of the method — more than a
dozen on 1.7.0 — are not reported as yours. A subclass or wrapper that replaces
the method entirely is recognised too, but not by looking for it. LiveKit's
automatic answer is emitted from inside AgentActivity._generate_reply()
(agent_activity.py:1559), synchronously, so that frame is on the stack every
single time the automatic answer is created. Its absence therefore rules the
automatic answer out, whatever the reply turns out to be — which is what makes a
decorated override safe: functools.wraps copies the name, so a wrapper frame
is still literally called wrapper, and no amount of name matching would find
it. A frame calling itself generate_reply on one of the recorded sessions is used
as a weaker second reading, for builds where that anchor cannot be resolved at
all — and attach() may be called more than once, so every session it was
given counts, not only the most recent one.
Where livekit.agents cannot be imported at all, none of this evidence exists
and there is also no LiveKit to have generated anything — so a reply that
cannot be traced to the recorded session's own call is reported as unreadable
rather than assumed to be automatic. The cost is a caveat on a compatible
framework that generates replies by itself; the alternative was attributing
those replies to whoever spoke last and saying nothing.
LiveKit's own calls are still
flagged, because not all of them answer speech already under way: the
beta/workflows and IVR helpers ask the caller a question before the caller has
said anything, and the transcript that follows is the answer to it rather than
the question it answered. The events are identical either way. The post-tool reply is placed on the turn that ran the tool, which is
where it almost always belongs, but a caller talking over the tool call could
have prompted it instead. Either way the flag is set on every span of that turn,
including the LLM span, so an interrupted reply that reported tokens still
carries the caveat.
The caveat is also counted, not only recorded. The dashboard's fleet view
reports coverage.inferred_reply_turns for the range you are looking at, and
says so on the page when it is non-zero. Without that the flag reached the
archive and stopped there: a range's latency, token and cost figures could move
because a reply had been placed by reading the call rather than by the events,
and nothing on the page would have mentioned it. The turn is not subtracted
from the turn count — the exchange happened and its measurements are real; what
is in doubt is which turn they belong to.
If you are running a build of livekit-agents that does not expose these
signals, no reply is merged and the flag is set — the recorder never guesses
silently because a field was missing. The same applies to the call stack: a
failure to read it is not treated as evidence that the reply was automatic.
Treat a turn carrying this flag as unreliable for response latency specifically. Everything else on it — audio, durations, token counts — is measured, not inferred.
reply_skipped separates a turn the agent chose not to answer from one whose
reply failed to sound. Both look identical in the numbers — a transcript, no TTS
span — and they call for opposite responses, so the absence of a reply is never
by itself reported as a declined turn. It is set only when the agent said so,
and requires agent= on observe_agent_session.
A caller's speech window ends at speech_ended, which the voice detector stamps
at the first pause. When one message arrives as several finals, the later ones
are evidence the caller kept talking, so the window ends at the last final less
the recogniser's own transcription_delay_ms, and presentation_window.source
reads final_transcripts to say so.
A session with events_complete: false or a non-zero drop count is classified
unverifiable in the dashboard, not healthy and not failed. "We do not know
what the caller heard" is a distinct claim, and it is ranked above merely slow
calls precisely because it hides problems.
events.jsonl
One JSON object per line, in append order. Operation spans are written when they
end and are identified by their type; audio and websocket lifecycle events
carry a kind field instead (audio_chunk, websocket, and — in Python —
capture_error).
operation
Written when the operation ends.
{
"event_id": "b41c…",
"session_id": "9f2c…",
"turn_id": "turn-2",
"scope": "turn",
"type": "llm",
"endpoint_id": "llm",
"provider": "openai",
"model": "gpt-4o-mini",
"transport": "http",
"started_at_ms": 4821,
"ended_at_ms": 6104,
"duration_ms": 1283,
"status": "ok",
"request": { "messages": "…" },
"response": { "tokens": 128 },
"error": null,
"milestones": {
"first_token": { "occurred_at_ms": 5290, "last_at_ms": 5290, "count": 1 }
},
"samples": {
"partial_transcript": { "items": [{ "occurred_at_ms": 5010, "text": "how do i" }], "truncated": false }
}
}| Field | Notes |
|---|---|
type | stt | llm | tts | tool |
scope | turn (default) or connection |
transport | manual, or set by instrumentation (http, websocket) |
status | ok by default; anything else marks the span failed |
milestones | Repeated names merge: first occurred_at_ms, latest last_at_ms, count |
samples | Per-name bucket, capped at limit (default 100), with a truncated flag |
audio_chunk
{ "kind": "audio_chunk", "track": "caller", "occurred_at_ms": 1240, "byte_length": 3200, "duration_ms": 100 }track is caller or agent. duration_ms is present when the PCM duration is
computable. The agent track's occurred_at_ms comes from the playout clock,
not arrival time — see below.
websocket
{ "kind": "websocket", "session_id": "9f2c…", "occurred_at_ms": 980, "event": "closed", "url": "wss://…", "code": 1000 }Payload bounding
Any captured payload larger than payload_max_bytes (16 KiB default) is
replaced in place:
{ "_truncated": true, "_original_bytes": 184320, "_preview": "{\"messages\":[{\"role\"…" }A value that cannot be serialised becomes { "_capture_error": "…" }. The event
is still written — capture never fails an operation.
call.audio
Raw interleaved stereo PCM. No container, no header.
| Property | Value |
|---|---|
| Encoding | pcm_s16le |
| Channels | 2 |
| Sample rate | The maximum of the two input track rates |
| Left channel | Agent |
| Right channel | Caller |
| Frame size | 4 bytes |
Bytes are frames × 4, where frames is the larger of the session duration in
samples and the longest rendered track. A track with no audio is silence.
The playout clock
TTS audio usually arrives in a burst even though it will be played in real time. If chunks were placed at arrival time, a 12-second reply would appear as a half-second blob and every pause in the call would vanish.
Instead, the agent track's clock advances by the PCM duration of each chunk:
a chunk is placed at max(arrival_time, end_of_previous_chunk). That preserves
the real silences, which is what makes the waveform's gap annotations meaningful
and what makes latency audible when you scrub to a timestamp.
Playing it
ffplay -f s16le -ar 16000 -ch_layout stereo call.audioffmpeg -f s16le -ar 16000 -ac 2 -i call.audio call.wavThe dashboard does the same thing on demand — it wraps the raw PCM in a WAV
header per request rather than storing a second copy, and honours HTTP Range
because Safari refuses to play media otherwise.
Audio dominates storage. Two raw 16-bit PCM tracks are roughly 64 KB per second of call at 16 kHz — about 3.8 MB per minute. Stored audio is raw PCM; transfer compression does not shrink stored objects. Upload methods do not remove spool copies, though the Python drainer has separate cleanup settings. The local dashboard has no automated retention policy. Plan disk capacity, backup and cleanup before increasing call volume.
Consuming a package yourself
The format is deliberately boring, so you do not have to use the dashboard:
import json
from pathlib import Path
package = Path("./.vaani-spool/9f2c1e4a-…")
manifest = json.loads((package / "manifest.json").read_text())
operations = [
json.loads(line)
for line in (package / "events.jsonl").read_text().splitlines()
if line and json.loads(line).get("type") in {"stt", "llm", "tts", "tool"}
]
turns = {}
for op in operations:
if op.get("scope") == "connection":
continue # a session-long socket is not an utterance
turns.setdefault(op["turn_id"], []).append(op)Two rules to honour if you build on this: skip scope: "connection" spans when
computing per-turn metrics, and treat a missing milestone as unmeasurable
rather than substituting an operation's start or end time. Both are the reason
the dashboard's numbers hold up.
Next
STT evaluation
Optional external comparison of recorded caller audio and production transcripts — model disagreement, streaming timing, semantic-risk judging and cost limitations.
HTTP API
Local dashboard API with explicitly marked hosted-development differences — upload, media, evaluation policy, pricing and workspace-deletion boundaries.