Vaanieval
Reference

Metrics glossary

Vaanieval timing and comparison metrics, their evidence requirements and interpretation limits.

Missing required evidence means unavailable, not zero. Keep measured timing separate from explicitly estimated costs, model disagreement and inferred span attribution. Everything below states what each number needs.

Latency

Reply wait

The caller-visible silence between finishing speaking and hearing a response.

Measured fromcaller stops speaking → first audio byte
Reported asTypical wait (median) and worst wait per call; p50/p95 across a fleet
Unavailable whenEither milestone is missing on the turn

This is not a provider's response time. A provider that answers in 200 ms can still produce a 3-second reply wait if endpointing was slow or a retry happened. Conflating the two is the most common mistake when moving from a request-level APM.

Time to first partial

Measured fromCaller speech start → first partial transcript
Tells youHow quickly the recognizer produced any text
Unavailable whenThe build did not record a first-partial milestone, or word timestamps are absent

Caller speech start is taken from word timestamps, not from speech_started — that milestone marks when the recognizer opened its listening window, which on real data can precede the first word by more than ten seconds.

Endpointing delay

Measured fromcaller stops speaking → provider marks speech final
Tells youHow long the recognizer took to decide the caller had finished
Unavailable whenThe two instants were not observed separately

Some STT spans stamp speech_ended and final_transcript from the same underlying framework event, so the milestones exist, look independent, and are byte-identical. Subtracting them would manufacture a zero, so the engine requires genuinely separate observations. This is why ENDPOINTING frequently shows — unavailable on real calls.

Finalization delay

Measured fromLast word → final transcript
Tells youHow long after the caller stopped the final text arrived
Unavailable whenWord timestamps or the final-transcript milestone are missing

Time to first token

Measured fromrequest sent → first token
Unavailable whenThe provider or framework does not report TTFT

Time to first byte (TTS)

Measured fromtext handed to the voice → first audio byte
Unavailable whenThe provider does not report TTFB

Coverage and quality of measurement

Measurable turns

The percentage of turns with every milestone a metric needs. Published alongside every percentile — "68% of 111 turns measurable".

A low measurable percentage is a finding about your instrumentation, not your agent. It usually means a provider wrapper is not emitting milestones or a code path bypasses your endpoint rules. Fix it before trusting the percentile above it.

Coverage

Recorded turns that carry usable STT evidence, e.g. "3 of 4 recorded turns".

Transcripts available

How many turns have transcript text at all, e.g. "3 of 3". Distinct from coverage: a call can have timing without text when stt_content capture is off.

The engine distinguishes a privacy-redacted build (records a character count but not the words) from one that never captured a transcript. "Withheld" and "absent" are different facts and are not conflated.

Reliability

Failures

Operations that ended with a status other than ok, or with a non-null error.

Cancellation is not failure. AbortError, CancelledError and CancelledException are recognised and excluded — a TTS span aborted because the caller barged in is the agent behaving correctly. The UI goes further and checks whether caller speech actually overlaps the cancellation, flagging one that was not barge-in as worth investigating.

Turns over 3s

Turns whose reply wait exceeded 3 seconds. A blunt but useful "how often did this feel slow" counter.

Attention classes

Calls are ranked in three classes, then by magnitude within each:

ClassMeaning
FailedA recorded operation failed; this alone does not prove a caller-visible failure
UnverifiableCapture is incomplete, so we cannot say what they heard
SlowA measured reply wait crossed a threshold; not a guarantee of other call quality

Unverifiable is ranked above slow deliberately. "We do not know what happened" is a different claim from "it was slow", and folding it into a healthy bucket would hide the calls most likely to contain a problem. Its sources are a partial session status, events_complete: false, or non-zero drop counts in capture_status.

Transcription disagreement

Estimated WER

Word error rate between the production transcript and a challenger transcript, from deterministic normalization and word alignment. Reported per turn and per call, with substitution, deletion and insertion counts.

This is a disagreement score, not accuracy. The challenger is a pseudo-reference, not ground truth — where the two models disagree, either could be wrong. The UI labels it estimated WER / model disagreement, never accuracy, and it stays that way until a human-reviewed reference transcript exists.

Unavailable when there is no usable challenger comparison. Running a comparison enables a disagreement score, not a ground-truth accuracy claim.

Semantic risk

An LLM judge estimates whether a disagreement could have changed the conversation. It can misclassify risk and is assistance for human review. It sends transcript comparison content to OpenAI; OPENAI_API_KEY is required. The model is set by STT_EVAL_JUDGE_MODEL (default gpt-4o-mini), with compatible OpenAI model fallback. A matching cached fingerprint can avoid another judge request, but reruns, cache misses and retries are not guaranteed free.

This exists precisely because WER alone over-reports: it counts every word equally, and conversations do not.

Cohort comparison

The same call's metrics against other calls from the same agent, computed over a 25-session sample (COHORT_SAMPLE_LIMIT) rather than full history — cheap, but a sample estimate.

Cost

Per-minute list pricing with a provenance flag:

ProvenanceMeaning
estimated_list_pricePublished list rate
configuredA rate you set via PUT /v1/pricing
usage_not_reportedThe provider did not report usage for this run

Cost figures are for comparing options, not reconciling invoices. Challenger transcription and judge requests may incur provider charges. The local path has no global spend cap; GET /v1/evaluation-policy returns cost_estimate and allowance as null, not zero.

Scoping rules that affect every number

1. Connection spans are excluded from per-turn statistics. A session-long STT websocket is a connection span, never scored as an utterance — otherwise a five-minute call would report a five-minute transcription latency.

2. Framework spans cover retries. An LLM operation whose duration exceeds any single HTTP attempt is timing the whole retry sequence. That is the caller's reality; the attempts are the mechanism.

3. Missing required milestones exclude a turn, they do not zero it. Operation duration or a batch round trip cannot substitute for missing endpointing or word-timing evidence.

Verification

scripts/validate-latency.py in the dashboard repository is a deliberate second implementation. It re-derives every published latency value straight from events.jsonl with its own arithmetic and asserts the served payload agrees, and it runs in the test suite — so a change that reintroduces a fabricated measurement fails the build.

Next

On this page