Self-hosting
Running the Vaanieval local dashboard — setup, configuration, external evaluation and operational limits, distinct from pending hosted beta work.
The dashboard is a FastAPI service that ingests session packages, serves the call console, and runs the STT evaluation engine.
Localhost only. Read the limits before using sensitive
recordings. Ingest authentication is available but off by default, read endpoints are
always open, and there is no tenant isolation and no retention policy. It is a
development and debugging tool, not a multi-tenant service. app.main:app
refuses hosted-environment startup unless explicitly configured for the curated
demo. Hosted-beta work uses the separate app.cloud.main:app entrypoint and is
not a public SaaS release or production-readiness certification.
Requirements
- Python 3.11 or newer
- Disk for audio (a 5-minute stereo call at 16 kHz is roughly 19 MB of raw PCM)
Install the runtime dependencies from this checkout's requirements.txt.
Source installs and hosted beta
The repositories are public and support anonymous Git access. These guides install the SDKs from public Git. Anonymous registry checks on 20 September 2026 returned HTTP 404 for npm @vaanieal/observer and PyPI vaanieval-observer; no release was found under those names. Source availability does not mean the hosted service is ready for public launch: new public workspaces remain closed.
The quickstarts pin published Python 0.5.7b1 and Node 0.1.1-beta.1 source previews. Both installed previews have uploaded synthetic calls to the isolated hosted development environment. Git/build tooling and explicit revision upgrades are required; this is not registry publication or production-readiness certification.
For setup help, email shubham@vaanieval.com or book a call — a short conversation about your use case tells us whether Vaanieval fits, and helps identify the setup that matches your stack.
Setup
Clone
git clone https://github.com/shubhamofbce/vaanieval-observer-backend.git dashboard
cd dashboardInstall
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install -r requirements-dev.txt # for testsRun
uvicorn app.main:app --host 127.0.0.1 --reload --port 8000- Console:
http://localhost:8000/ - Sessions list:
http://localhost:8000/dashboard - Alert previews (browser-local, no delivery):
http://localhost:8000/alerts - Setup checklist:
http://localhost:8000/onboarding - STT evaluation workspace:
http://localhost:8000/stt-evaluation?session=<id> - Health:
http://localhost:8000/health
Point an SDK at it
VaaniObserver(endpoint="http://localhost:8000", api_key="local-dev")The SDK requires a non-empty key, but local ingest does not enforce key validity by default. Recognised keys can record usage for onboarding. See Authentication to enable the ingest gate.
Authentication
Ingest enforcement is off by default. Turning it on does not make this local service safe to expose beyond localhost.
| Variable | Default | Effect |
|---|---|---|
VAANI_REQUIRE_API_KEY | 0 | 1 requires a bearer token on the two ingest endpoints |
With the gate off, ingest accepts requests without a valid key. Usage of a
recognised key can be recorded for onboarding. With it on, POST /v1/sessions
and POST /v1/sessions/{id}/complete return 401 without a live key.
Keys are managed over HTTP:
# Mint the first key. While enforcement is on and no active key exists, this is
# allowed from loopback only — which is why you run it on the host itself.
curl -sX POST http://localhost:8000/v1/api-keys \
-H 'content-type: application/json' -d '{"name":"agent-fleet"}'
curl -s http://localhost:8000/v1/api-keys # list (never returns secrets)
curl -sX DELETE http://localhost:8000/v1/api-keys/<id> -H 'authorization: Bearer <key>'Then set VAANI_API_KEY in the agent's environment to the returned token.
The gate covers writes only. Every read path — the console, GET /v1/sessions, and GET /v1/sessions/{id}/audio/{track} — stays open with
enforcement on. Local object PUTs are also unauthenticated; they are not
signed object-store URLs. An SDK must not forward the ingest bearer token to
arbitrary upload URLs. Treat VAANI_REQUIRE_API_KEY as a limited ingest gate,
not read authentication, tenant isolation or permission to host this API.
Configuration
| Variable | Default | Purpose |
|---|---|---|
VAANI_DATA_DIR | ./data | Runtime data directory |
VAANI_REQUIRE_API_KEY | 0 | 1 requires a bearer token on ingest — see Authentication |
VAANI_EVALUATIONS_ENABLED | off | Only 1 opts into manual external evaluation; provider keys alone do not enable it |
ELEVENLABS_API_KEY | — | Challenger transcription; required to run an STT evaluation |
OPENAI_API_KEY | — | Semantic risk judge |
STT_EVAL_JUDGE_MODEL | gpt-4o-mini | OpenAI judge model; a compatible OpenAI model may be used as fallback |
VAANI_ENV_FILE | ./.env | os.pathsep-separated dotenv paths to read keys from when they are not exported |
Never commit provider keys. .env and data/ are both gitignored. Note that
ELEVENLABS_API_KEY and OPENAI_API_KEY are billed per use — a batch of
challenger evaluations against long calls can cost real money. The local path
has no global spend cap or allowance enforcement.
Before running a comparison
Fetch GET /v1/evaluation-policy and read its enabled, reason, disclosure
and consent_version. After opting in, Run comparison requires confirmation
of the current disclosure; model selection alone does not run a job. API callers
submit the current consent_version with model: "elevenlabs_scribe_v2".
The full recorded caller track is sent to ElevenLabs Scribe v2, and transcript
comparison content is sent to OpenAI for semantic-risk judging when configured.
The policy reports cost_estimate: null and allowance: null: those are unknown,
not zero or unlimited credits. See STT evaluation
and Capture and privacy.
Storage layout
Everything lives under VAANI_DATA_DIR:
Relevant local SQLite tables include:
| Table | Contents |
|---|---|
sessions | id, manifest_json, status, created_at, updated_at, completed_at |
operations | id, session_id, operation_json, started_at_ms, turn_id, scope, failed |
challenger_evaluation_jobs | session_id, model_key, job_id, status, error, timestamps |
Audio is stored once, as the SDK's raw PCM. The console requests an on-demand
WAV wrapper for browser playback rather than storing a second copy. The wrapper
streams and honours HTTP Range, which Safari requires before it will play any
media response.
PRAGMA user_version guards a one-time backfill of the failed column on
upgrade.
Ingestion
Three steps, described in full in the HTTP API reference:
POST /v1/sessions
The manifest, with an idempotency-key header equal to the session id. A
mismatch is a 400. Returns upload URLs.
PUT /v1/uploads/{session_id}/{object_name}
Only four names are accepted: events.jsonl, call.audio, caller.audio,
agent.audio. Bodies stream to a .part file and are renamed on success. The
cap is 128 MiB; beyond that the server returns 413.
POST /v1/sessions/{session_id}/complete
Byte size and SHA-256 for each object. Both are verified; a mismatch is a 400.
The session becomes ready if operations were imported, or partial if not.
partial is what surfaces as unverifiable in the dashboard. It means the
objects arrived but no operations could be read — a crash before
session.end(), capture disabled mid-call, or an agent that never spoke.
Those need opposite responses, so the console reads
capture_status.measured.agent_audio_ms where available. Interpret a zero
only alongside whether the audio tap was installed and capture was complete;
missing capture is not proof that the agent was mute.
Restart behaviour
Challenger evaluation jobs that were queued or in_progress when the process
stopped are marked failed on startup. They are not resumed. These are local
in-process jobs, not a durable cloud queue. Re-queue from the UI after reviewing
the disclosure again; a failed job may already have incurred provider charges.
Tests
pytestTests run against a temporary data directory and SQLite file per test, so they
never touch data/.
scripts/validate-latency.py is a deliberate second implementation: it
re-derives every published latency value straight from events.jsonl using its
own arithmetic and asserts the payload agrees. It is wired into the suite, so a
change that reintroduces a fabricated measurement fails the build. It needs
recorded calls in data/, so it skips on a clean checkout.
Operational limits
These are real constraints, not future work items. Decide against them before you depend on the dashboard.
Security
- Ingest authentication is available but off by default. Set
VAANI_REQUIRE_API_KEY=1and mint a key (POST /v1/api-keys) to require a bearer token onPOST /v1/sessionsandPOST /v1/sessions/{id}/complete. Left off, any non-empty key is accepted and recorded but not checked. - Read endpoints are never authenticated. The console, the session API and
GET /v1/sessions/{id}/audio/{track}stream a recorded call to any caller who can reach the service and knows the session id — with or withoutVAANI_REQUIRE_API_KEY. - No tenant isolation. Every session is visible to everyone who can reach the service.
Enabling the key gate stops anonymous writes; it does not make the dashboard safe to expose. Bind it to localhost. Neither an ingest key nor a reverse proxy adds tenant isolation to this application.
Scale
| Component | Limit | What happens past it |
|---|---|---|
| SQLite metadata | Single file, single writer | Concurrent ingestion serialises; write contention grows with volume |
| Audio storage | Local filesystem | No replication; the host's disk is the ceiling |
| Upload size | 128 MiB per object, decompressed | 413; a very long call cannot be uploaded. Content-Encoding: gzip reduces transfer time, not the cap |
| Challenger evaluation | 2-worker thread pool, in-process | Replays compete with request serving; a queue of long calls makes the console sluggish |
| Cohort comparison | 25-session sample (COHORT_SAMPLE_LIMIT) | Comparisons are computed against a sample, not the full history |
Multi-instance operation over the same local data directory is not supported. SQLite write contention and shared-file races grow with concurrency; use one local instance rather than treating shared storage as a scaling strategy.
Data lifecycle
- No retention policy. Sessions and audio accumulate until you delete them.
- No deletion API. Plan manual cleanup of related rows, objects, evaluation artifacts and backups; this is not a legal erasure guarantee.
- No backup.
vaani.dbandobjects/are ordinary files; back them up yourself. - SDK spool directories are not cleaned up by
upload_package()either, so audio accumulates on your agent hosts as well. Runpython -m vaani_observer.drain --watch 300there.
Operations
- Single process, no queue: an ingestion spike is absorbed by the web workers.
- No metrics endpoint and no structured audit log.
- Schema migrations are a single
user_versionguard, not a migration framework.
When to use it anyway
Use it for a bounded localhost debugging workflow, preferably starting with synthetic calls. The simple setup trades managed access, durable jobs, automated retention and scaling for local control; those gaps matter even at low volume when recordings are sensitive.
Hosted-beta development is separate. Do not use the curated demo as an ingest
endpoint or expose app.main:app behind a key as a tenant-safe service. Fleet
rollups and browser-local alert previews do not imply notification delivery or
hosted production readiness.
Hosted development boundary
The checkout's hosted API is app.cloud.main:app; it uses a separate worker
(python -m app.cloud.worker) and PostgreSQL, rather than the local SQLite setup
above. Both explicit migrations in app/cloud/migrations/ are required:
001_initial.sql, then 002_media_and_workspace_deletion.sql. Startup requires
schema version 2.
Cloud dependency installs, including the container build, use:
pip install -r requirements-cloud.txt -c requirements-cloud.lockThe Docker base image is pinned by digest. The build context includes the cloud
runtime and hosted web assets but excludes the legacy local application and
static console. Pinning improves reproducibility at the cost of explicit
maintenance: refresh the constraints and base digest for security updates, then
rerun tests and Linux AMD64 image validation. A successful non-root, read-only,
no-network import and pip check smoke test does not establish live cloud
permissions, recovery or soak-test readiness.
Admission defaults closed, with explicit controlled-beta enablement, a subject allowlist and open admissions required.
The hosted working tree includes separate WAV playback/media-link and
asynchronous workspace-data deletion paths, not a public availability guarantee.
Global playback support does not mean a recording is prepared: check that
recording's preview.available. Migration 002 does not silently re-render older
calls.
Workspace deletion requires recent factor verification from signed Clerk fva
claims and confirmation of the authenticated workspace. It does not delete the
Clerk identity. The UI confirmation is DELETE WORKSPACE; the API confirmation
value is DELETE, with the authenticated workspace ID. A minimal ledger remains:
workspace identifier, hashed owner subject, deletion timestamps and monthly usage
totals, not call content or API
keys. This is not a legal erasure guarantee. Derived turns, analytics
and evaluations remain unavailable; no evaluation providers are called. Tests
do not prove runtime managed-identity permissions, restore/rollback or policy/privacy
readiness. This page remains a localhost setup guide, not authorization to
launch the hosted service publicly.
An isolated owner-beta API, always-on worker, private PostgreSQL 17 database, private Blob container, Key Vault and separate managed identities are deployed in Azure Central India at beta.vaanieval.com, with valid managed HTTPS. Public workspace admissions remain closed. Migrations 1 and 2 and the live runtime database privilege gate passed. Synthetic checks against temporary private Azure Blob storage passed independently using operator Entra CLI credentials. They covered bounded ranges, ETags, immutable writes, private-storage policy, uncommitted-block cleanup and purge. A mixed in-process API, real PostgreSQL and real Blob journey also passed upload, WAV preparation, playback and workspace purge with simulated application identity. The temporary test storage and access grant were removed afterward.
Those earlier storage-only checks do not establish current runtime ingestion,
recovery or soak acceptance. A production Clerk instance was created through the
existing browser session on 20 September 2026, preserving Google-only sign-in.
The five Clerk CNAMEs are now installed through Spaceship; Clerk reports DNS,
TLS, email and Google OAuth complete. The original 18 DNS records and their TTLs
are unchanged. A real Google consent/callback created a verified Google-linked
production owner account, and https://accounts.vaanieval.com/user renders its
signed-in profile. No passwords were handled or authentication bypassed.
The dedicated Google OAuth project is vaanieval, with only OpenID, email and
profile scopes. Its audience is now External / In production, with the
deployed beta homepage, /privacy and /terms links saved. Google verified the
Vaanieval branding and it was published; a reloaded console confirms it is being
shown to users. Google's
basic-sign-in exception
exempts these scopes from tester-list and seven-day authorization restrictions;
that is not brand verification. The hosted policy pages describe this beta's
actual data flow and limits. They are draft beta policies, not a legal-review
claim; the unrelated self-hosted website policy was not reused.
The first completed login exposed a fallback to the then-undeployed beta hostname.
Account Portal sign-up/sign-in/logo fallbacks now persist as
https://beta.vaanieval.com, confirmed by reloading the production settings.
The earlier account-profile fallback is no longer the application destination.
The real owner now has a workspace created through the hosted UI. Admission was
briefly enabled for that existing allowlisted subject only and then closed again.
Both installed published SDK previews uploaded synthetic span-only and silent-audio
calls; all four reached ready through the deployed worker. Owner call/span review,
ingest-key read isolation and post-revocation rejection passed. Deleting one
explicitly disposable call removed its events object from primary Blob storage;
the seven other objects were unchanged. The metadata audit excluded snapshots,
versions and soft-deleted copies, and its temporary reader grant was removed.
Privileged bootstrap access is retired and the administrator credential was
rotated and proved inside the private network; temporary proof access was removed.
This is not second-customer isolation, restore, fresh-sign-in workspace deletion
or broad device/playback acceptance. Public launch remains no-go.
Independent UX findings led to explicit no-audio/unknown/pending recording messages and readable mobile table columns inside a labelled horizontal-scroll region. Both fixes are deployed.
The owner environment now keeps one warm API replica and one always-on worker, with a small database without HA. Actual scale-from-zero events correlated with repeated request timeouts, so API scale-to-zero was removed. This increases idle cost: revisit the original $100-120 estimate with $120-140/month before-tax planning headroom, not a quote or cap. Single-region downtime and limited throughput remain constraints; a warm singleton is not HA or a latency SLA. Never stop the retention worker while deletion obligations exist.
The current Azure subscription is a confirmed Visual Studio/MSDN offer with its spending limit enabled. Microsoft's monthly-credit terms limit it to development/testing and reserve suspension rights. This environment is for owner-only synthetic acceptance, not customer production. An eligible production subscription is required before customer use; billing settings were not changed.
Five basic resource-scoped metric alerts are deployed. Operator email verification remains pending, so notification delivery is not certified.
The two Microsoft-owned service databases retain CONNECT/TEMP and exact read-only Query Store privileges. SQL/wait capture, text emission, plans and utility capture are disabled, and previous Query Store history was reset. Application/default-database boundaries and the absence of service writes, grant options, ownership and security-definer access were checked using the actual API/worker roles. This is not exclusive database connectivity; adding a second application or re-enabling query capture requires another isolation review.