Vaanieval
Concepts

Capture and privacy

Vaanieval capture settings, local storage, explicit dashboard upload and optional external evaluation — including data exposure and operational limits.

Voice agents handle some of the most sensitive data a company holds: recorded speech, transcripts of that speech, and the prompts built from it. Vaanieval captures only what your settings and integration supply. The core SDK and LiveKit integration have different defaults; review both before recording.

Defaults

VaaniObserver(
    capture={
        "audio": True,
        "http_bodies": False,
        "websocket_text_frames": False,
        "stt_content": False,
        "payload_max_bytes": 16384,
    },
)
FlagDefaultWhat it controls
audioonRecording caller and agent PCM
http_bodiesoffRequest and response bodies on instrumented HTTP calls
websocket_text_framesoffCapture configuration flag; websocket observation described here records lifecycle and counts, not frame contents
stt_contentoffTranscript text inside STT operations
payload_max_bytes16384Size bound on every captured payload

Core audio capture is on by default when audio frames are supplied. Calls can contain identifying or sensitive information. Review recording permission, access, retention and deletion requirements for your use case; these docs provide no legal or compliance guarantee. Disable audio to avoid recordings, but expect audio-dependent review and metrics to be unavailable.

LiveKit differs: VAANI_ENABLED defaults to false, but once enabled its VAANI_CAPTURE_HTTP_BODIES, VAANI_CAPTURE_STT_CONTENT and VAANI_UPLOAD defaults are true. Set these explicitly rather than relying on core defaults. See LiveKit Agents.

Instrumentation boundaries

The built-in transport observers are not general-purpose packet capture:

  • Authentication headers are not collected by the built-in HTTP and websocket observers. This is not a credential-redaction guarantee: URLs, metadata, manually supplied payloads and enabled body capture can contain secrets or identifying information.
  • Streaming response bodies are not drained for capture. Doing so would hold back your first token.
  • Frame-by-frame websocket contents are not recorded by socket observation. It records byte counts, frame counts and timings.
  • Automatic provider capture requires a matching endpoint rule and ambient session (or an explicitly forced endpoint). Manually recorded operations and integration hooks can still write data independently of those rules.

Size bounding

Every captured payload — request, response and sample data — passes through a bound of payload_max_bytes (16 KiB by default). Oversized values are replaced:

{ "_truncated": true, "_original_bytes": 184320, "_preview": "{\"messages\":[{\"role\"…" }

Values that cannot be serialised become {"_capture_error": "…"}. Bounding is not redaction: a retained preview can still contain sensitive content.

Truncation is visible, never silent. If you see _truncated in a trace, the payload was clipped. Check timing and capture coverage separately; truncation alone does not establish measurement correctness.

Turning capture up

Debugging a specific problem usually justifies temporarily capturing more. Do it narrowly:

vaani = VaaniObserver(
    capture={
        "http_bodies": True,       # see the exact prompt sent
        "stt_content": True,       # see the transcript the model received
        "payload_max_bytes": 65536,
    },
)

stt_content: True means caller speech in plain text lands in events.jsonl and in the dashboard's SQLite database. http_bodies: True means prompts — which typically embed that transcript plus retrieved customer records — can land there too. Review the data you permit before enabling these flags. Prefer synthetic calls when debugging capture.

Note that STT evaluation requires transcript content. You cannot compute an estimated WER against a challenger model without the production transcript to compare. If you want the STT evaluation workspace, stt_content must be on for the calls you intend to review.

Where the data goes

Local disk, during the call

Captured data is written under spool_directory (default ./.vaani-spool). Capture does not itself upload to the dashboard during the call. Your agent's normal provider traffic is separate and may already involve external services.

Your dashboard, after the call

Upload is an explicit step via upload_package() / uploadPackage(), or a configured integration/drainer. The LiveKit recorder uploads after finish() unless VAANI_UPLOAD=false or upload=False. The configured dashboard receives the captured events and audio; local spool copies are not automatically removed by the upload method.

Optional external evaluation, after upload

The local dashboard defaults evaluations off. Enabling VAANI_EVALUATIONS_ENABLED=1 allows a manual, confirmed comparison:

  • ElevenLabs Scribe v2 (scribe_v2) receives the full recorded caller audio for the session, not just the selected turn or visible clip.
  • OpenAI receives transcript comparison content for semantic-risk judging when configured. STT_EVAL_JUDGE_MODEL selects the model; a compatible OpenAI model may be used as a fallback.

Read GET /v1/evaluation-policy for the current disclosure and consent_version. Choosing a model does not execute a job. The UI's Run comparison action requires confirmation; API callers must submit that current version with the request. This records acknowledgement of the processing disclosure, not legal permission to process somebody else's data.

Provider processing may incur charges. The local policy supplies no run-cost estimate or allowance (null), and there is no global spend cap in this path. Do not assume anonymity, zero retention or provider data deletion from these settings. Review the applicable provider terms and your account configuration before sending data.

Failure behaviour, and why it matters here

By default strict is False. Instrumentation errors are swallowed so that these failures usually degrade capture rather than propagate.

VaaniObserver(strict=True)   # raise instead of degrading, for staging

The cost is that capture can silently degrade. Two things make that visible:

capture_status in the manifest:

{
  "capture_status": {
    "events_complete": true,
    "audio_complete": false,
    "http_instrumentation": "active",
    "websocket_instrumentation": "active",
    "dropped_event_count": 0,
    "dropped_audio_chunk_count": 37
  }
}

The unverifiable class in the dashboard. A call whose capture is incomplete is ranked separately from failed and slow calls, because "we do not know what the caller heard" is a different statement from "the caller heard an error". It is never quietly folded into the healthy bucket.

Retention and deletion

Stated plainly, because the gap matters:

The local dashboard has no automated retention policy, deletion API or authenticated reads. Sessions, audio and evaluation artifacts accumulate until you manage them. Optional ingest keys do not make this tenant-safe. Keep it on localhost and plan cleanup across database rows, files, SDK spools and backups. Local deletion does not establish deletion by external providers. See Self-hosting.

Spool directories are also not cleaned up automatically after a successful upload. On a busy agent host this is a disk-space concern; add a sweep of .vaani-spool for completed sessions to your deployment.

Capture choices and tradeoffs

SettingSensitive callsSynthetic debugging calls
audioOff unless recording is approved; no audio review when offOn
http_bodiesOffOn
stt_contentOff, unless you need STT evaluationOn
websocket_text_framesOffDo not rely on it for frame-content capture
payload_max_bytes1638465536
strictFalseTrue
UploadExplicit, post-callExplicit, post-call

Less capture reduces data exposure and disk use but also reduces review coverage. More calls increase spool, dashboard and backup storage; optional evaluation adds external processing, provider charges and in-process queue contention. These settings do not turn the local dashboard into a hosted production service.

Next

On this page