Vaanieval
Dashboard

STT evaluation

Optional external comparison of recorded caller audio and production transcripts — model disagreement, streaming timing, semantic-risk judging and cost limitations.

Latency helps investigate responsiveness. STT evaluation compares transcripts to identify disagreements worth listening to; it does not establish what was said or prove that either model heard correctly.

STT review for an isolated localhost synthetic tool-only call: coverage 0 of 1, transcripts 0 of 0, all speech timing unavailable and external evaluations disabled.

This synthetic fixture contains no conversation or transcription evidence, so the empty metrics are expected. The screenshot is not a hosted customer call, and no evaluation-provider request was made to create it.

The two questions it answers

  1. Where did production STT disagree with a challenger transcript, and should I listen to the audio?
  2. Did the production model detect and finalize caller turns quickly and reliably enough for a live agent?

Open it at /stt-evaluation?session=<id>, or from the STT review tab on a call.

Streaming timing

Available immediately, from milestones captured during the call:

MetricMeasured from
First partialCaller speech start → first partial transcript
FinalizationLast word → final transcript
EndpointingCaller stops speaking → provider marks speech final
CoverageRecorded turns with usable STT evidence

— means unavailable, and it always carries a reason. In the screenshot, ENDPOINTING is — unavailable. It is never rendered as 0, never inferred from an operation's start and end times, and never derived from a batch HTTP round trip. If the milestones that define a metric were not captured, the metric does not exist for that call.

The engine is unusually careful about what counts as evidence:

  • Two identical timestamps may not be independent measurements. Some STT spans stamp speech_ended and final_transcript from the same underlying framework event, so the milestones can exist, look independent, and be byte-identical. Subtracting them would manufacture a zero. The engine requires the two instants to have been observed separately.
  • speech_started is a listening window, not speech. It marks when the recognizer opened its window, which on real data can precede the first word by more than ten seconds. Word timestamps are the only evidence of when the caller actually spoke.
  • Capture capability is tracked per session. A privacy-redacted build that records only a character count is distinguished from a build that never captured a transcript at all — so "no transcript" and "transcript withheld" are not conflated.

Comparison is optional external processing

The local dashboard defaults evaluations off. Only VAANI_EVALUATIONS_ENABLED=1 opts in; configuring provider keys is not enough. A run sends the full recorded caller audio to ElevenLabs Scribe v2 (scribe_v2), not just a selected turn or clip. Transcript comparison content is sent to OpenAI for semantic-risk judging when configured, potentially using a compatible OpenAI model fallback. Review data permissions and provider terms; no anonymity or zero-retention guarantee is made.

Choosing a model in the UI does not execute a job. Select Run comparison, review the current processing disclosure, and confirm.

For an API client, first call:

GET /v1/evaluation-policy

Read enabled, reason, challenger, semantic_risk, disclosure and consent_version. After acknowledging the current disclosure, submit its consent_version value (not the placeholder below):

POST /v1/sessions/{session_id}/challenger-evaluation
Content-Type: application/json

{ "model": "elevenlabs_scribe_v2", "consent_version": "<current policy value>" }

Poll the same path with GET. Jobs run off the request thread on a two-worker in-process pool, with status tracked in SQLite. The API returns 503 when evaluation is disabled, 409 for missing or stale consent, and 422 for an unsupported model key. Transcription requires ELEVENLABS_API_KEY.

The policy's cost_estimate and allowance are null in the local path. They do not represent a free run, a reserved budget or prepaid credits.

CHALLENGER_MODELS currently contains one entry: ElevenLabs Scribe v2.

What a run produces

Word alignment

Deterministic normalization and alignment of the production and challenger transcripts, producing substitution, deletion and insertion counts.

Estimated WER

Per turn and for the call, computed deterministically from stored production and challenger transcripts. The challenger transcript is a provider output, not ground truth.

Word-to-turn mapping

Each disagreement is attributed to the turn it occurred in, so you can jump straight to the audio.

A semantic risk judge

An LLM pass that estimates whether a disagreement could have changed the conversation. It helps prioritise human review, but can miss or misclassify risk. Requires OPENAI_API_KEY; the model defaults to gpt-4o-mini via STT_EVAL_JUDGE_MODEL, with compatible-model fallback within OpenAI. Cached results can avoid repeated judge requests when the fingerprint matches; cache misses, retries and new challenger runs can still incur charges.

A cost model

Per-minute list pricing with a provenance flag, batch→streaming substitution for switch decisions, and a monthly savings estimate.

The three interpretation rules

These are enforced throughout the engine, not just in the copy.

1. The challenger is a pseudo-reference, not ground truth. The score is labelled estimated WER / model disagreement, not ground-truth accuracy. Where the two disagree, either could be wrong. It is a disagreement score until a human-reviewed reference transcript exists — which is why the risk judge is review assistance, not a replacement for listening.

2. Unavailable is shown as — with a reason. Never 0, never inferred from an operation's start and end times, never derived from a batch round trip. A challenger that was only run as a batch transcription genuinely has no streaming milestones, and the UI says so rather than reporting zeros.

3. A session-long STT websocket is a connection span, not an utterance. It is never scored as a turn. Otherwise a five-minute call would report a five-minute transcription latency.

Cost

Evaluation can incur provider charges. Challenger transcription and judge requests are billable under your provider accounts. The local service has no global spend cap or allowance enforcement. Published cost figures are list prices with provenance (estimated_list_price versus configured), not invoices — they are for comparing options, not for reconciling a bill. Override the rates at PUT /v1/pricing.

The cohort comparison is computed against a 25-session sample (COHORT_SAMPLE_LIMIT), not your full history, which keeps a comparison cheap but means it is a sample estimate.

Verification

scripts/validate-latency.py is a deliberate second implementation. It re-derives every published latency value straight from events.jsonl with its own arithmetic and asserts the served payload agrees. It is wired into the test suite, so a change that reintroduces a fabricated measurement fails the build.

That is the strongest statement the project makes about its own numbers, and it is worth knowing it exists before you act on one.

Limits

  • One challenger model is currently supported.
  • Post-call, default-off and manual. Nothing is evaluated automatically on ingest; current-policy confirmation is required.
  • Requires transcript capture. stt_content must have been enabled for the call. See Capture and privacy.
  • Requires caller audio. The challenger replays the recorded caller track; a call captured without audio cannot be evaluated.
  • Two workers, in-process, not durable. Long calls compete with request serving and provider rate limits. Interrupted jobs fail on restart, not resume, and may already have incurred provider charges.

Next

On this page