STT evaluation
Optional external comparison of recorded caller audio and production transcripts — model disagreement, streaming timing, semantic-risk judging and cost limitations.
Latency helps investigate responsiveness. STT evaluation compares transcripts to identify disagreements worth listening to; it does not establish what was said or prove that either model heard correctly.

This synthetic fixture contains no conversation or transcription evidence, so the empty metrics are expected. The screenshot is not a hosted customer call, and no evaluation-provider request was made to create it.
The two questions it answers
- Where did production STT disagree with a challenger transcript, and should I listen to the audio?
- Did the production model detect and finalize caller turns quickly and reliably enough for a live agent?
Open it at /stt-evaluation?session=<id>, or from the STT review tab on a
call.
Streaming timing
Available immediately, from milestones captured during the call:
| Metric | Measured from |
|---|---|
| First partial | Caller speech start → first partial transcript |
| Finalization | Last word → final transcript |
| Endpointing | Caller stops speaking → provider marks speech final |
| Coverage | Recorded turns with usable STT evidence |
— means unavailable, and it always carries a reason. In the screenshot,
ENDPOINTING is — unavailable. It is never rendered as 0, never inferred
from an operation's start and end times, and never derived from a batch HTTP
round trip. If the milestones that define a metric were not captured, the metric
does not exist for that call.
The engine is unusually careful about what counts as evidence:
- Two identical timestamps may not be independent measurements. Some STT spans
stamp
speech_endedandfinal_transcriptfrom the same underlying framework event, so the milestones can exist, look independent, and be byte-identical. Subtracting them would manufacture a zero. The engine requires the two instants to have been observed separately. speech_startedis a listening window, not speech. It marks when the recognizer opened its window, which on real data can precede the first word by more than ten seconds. Word timestamps are the only evidence of when the caller actually spoke.- Capture capability is tracked per session. A privacy-redacted build that records only a character count is distinguished from a build that never captured a transcript at all — so "no transcript" and "transcript withheld" are not conflated.
Comparison is optional external processing
The local dashboard defaults evaluations off. Only
VAANI_EVALUATIONS_ENABLED=1 opts in; configuring provider keys is not enough.
A run sends the full recorded caller audio to ElevenLabs Scribe v2
(scribe_v2), not just a selected turn or clip. Transcript comparison content
is sent to OpenAI for semantic-risk judging when configured, potentially
using a compatible OpenAI model fallback. Review data permissions and provider
terms; no anonymity or zero-retention guarantee is made.
Choosing a model in the UI does not execute a job. Select Run comparison, review the current processing disclosure, and confirm.
For an API client, first call:
GET /v1/evaluation-policyRead enabled, reason, challenger, semantic_risk, disclosure and
consent_version. After acknowledging the current disclosure, submit its
consent_version value (not the placeholder below):
POST /v1/sessions/{session_id}/challenger-evaluation
Content-Type: application/json
{ "model": "elevenlabs_scribe_v2", "consent_version": "<current policy value>" }Poll the same path with GET. Jobs run off the request thread on a two-worker
in-process pool, with status tracked in SQLite. The API returns 503 when
evaluation is disabled, 409 for missing or stale consent, and 422 for an
unsupported model key. Transcription requires ELEVENLABS_API_KEY.
The policy's cost_estimate and allowance are null in the local path. They
do not represent a free run, a reserved budget or prepaid credits.
CHALLENGER_MODELS currently contains one entry: ElevenLabs Scribe v2.
What a run produces
Word alignment
Deterministic normalization and alignment of the production and challenger transcripts, producing substitution, deletion and insertion counts.
Estimated WER
Per turn and for the call, computed deterministically from stored production and challenger transcripts. The challenger transcript is a provider output, not ground truth.
Word-to-turn mapping
Each disagreement is attributed to the turn it occurred in, so you can jump straight to the audio.
A semantic risk judge
An LLM pass that estimates whether a disagreement could have changed the
conversation. It helps prioritise human review, but can miss or misclassify
risk. Requires OPENAI_API_KEY; the model defaults to gpt-4o-mini via
STT_EVAL_JUDGE_MODEL, with compatible-model fallback within OpenAI.
Cached results can avoid repeated judge requests when the fingerprint matches;
cache misses, retries and new challenger runs can still incur charges.
A cost model
Per-minute list pricing with a provenance flag, batch→streaming substitution for switch decisions, and a monthly savings estimate.
The three interpretation rules
These are enforced throughout the engine, not just in the copy.
1. The challenger is a pseudo-reference, not ground truth. The score is labelled estimated WER / model disagreement, not ground-truth accuracy. Where the two disagree, either could be wrong. It is a disagreement score until a human-reviewed reference transcript exists — which is why the risk judge is review assistance, not a replacement for listening.
2. Unavailable is shown as — with a reason. Never 0, never inferred from
an operation's start and end times, never derived from a batch round trip. A
challenger that was only run as a batch transcription genuinely has no streaming
milestones, and the UI says so rather than reporting zeros.
3. A session-long STT websocket is a connection span, not an utterance. It is never scored as a turn. Otherwise a five-minute call would report a five-minute transcription latency.
Cost
Evaluation can incur provider charges. Challenger transcription and judge
requests are billable under your provider accounts. The local service has no
global spend cap or allowance enforcement. Published cost
figures are list prices with provenance (estimated_list_price versus
configured), not invoices — they are for comparing options, not for
reconciling a bill. Override the rates at PUT /v1/pricing.
The cohort comparison is computed against a 25-session sample
(COHORT_SAMPLE_LIMIT), not your full history, which keeps a comparison cheap
but means it is a sample estimate.
Verification
scripts/validate-latency.py is a deliberate second implementation. It
re-derives every published latency value straight from events.jsonl with its
own arithmetic and asserts the served payload agrees. It is wired into the test
suite, so a change that reintroduces a fabricated measurement fails the build.
That is the strongest statement the project makes about its own numbers, and it is worth knowing it exists before you act on one.
Limits
- One challenger model is currently supported.
- Post-call, default-off and manual. Nothing is evaluated automatically on ingest; current-policy confirmation is required.
- Requires transcript capture.
stt_contentmust have been enabled for the call. See Capture and privacy. - Requires caller audio. The challenger replays the recorded caller track; a call captured without audio cannot be evaluated.
- Two workers, in-process, not durable. Long calls compete with request serving and provider rate limits. Interrupted jobs fail on restart, not resume, and may already have incurred provider charges.