NexaVoxa

A voice AI platform: build, test and run phone-answering agents with knowledge bases, numbers, campaigns and usage billing in one workspace.

4
Agents on the demo workspace
Browser
Test calls
Multi-tenant
Workspaces

What they came with

Phones still ring in businesses that have nobody spare to answer them, and the calls that do get through land on someone reading the same answers off the same page: opening hours, availability, price, can I book. Building a voice agent to take that work has meant stitching a telephony vendor to a speech model to an LLM to a billing system, then finding out about the latency in production. NexaVoxa is that stitching done once: an agent, a knowledge base, a phone number and a live test call in one workspace, with the minutes metered while the call is still running.

What the engagement covered

  • Single-prompt or guided-flow agents with model selection per agent
  • Test the agent from the browser before connecting a number
  • Knowledge bases, phone numbers, integrations and campaigns in one workspace
  • Role templates for receptionist, booking, outbound, lead qualification and support
  • Describe it in a sentence and let the platform draft the agent
  • Every setting stays editable after creation

Technical detail

The call runs as a stream

Twilio opens a bidirectional media stream over a WebSocket into FastAPI; Deepgram transcribes it live and the GPT-4o token stream is cut at sentence boundaries and handed to ElevenLabs, so the first audio goes back before the reply is finished. Incoming speech during playback cancels the queued TTS rather than talking over the caller.

Retrieval on a latency budget

Knowledge base documents are chunked and embedded into pgvector, and each turn retrieves against a hard time limit — if nothing comes back in time the agent says it does not know instead of stalling or inventing. Guardrails are prompt-level and per agent, so a receptionist template cannot quote a price it was never given.

Minutes metered on the call leg

Usage is written per second against the call record and drawn from the customer's credit balance, which is checked before the call is bridged so an empty balance refuses new calls rather than accruing debt. Stripe handles top-ups and invoicing, and every meter is reconciled nightly against Twilio's own duration record.

Draft agents and published agents

Agent configuration — prompt, voice, tools, knowledge base version — is versioned, and the live test call runs the whole pipeline against the draft. Publishing is an explicit step, so a number in production never picks up a half-edited prompt.

The stack

Interface

Next.jsReactTypeScriptTailwind CSS

Voice pipeline

TwilioDeepgramElevenLabsGPT-4o

Services and data

FastAPIPythonPostgreSQLpgvectorRedisCelery

Billing and hosting

StripeDockerTraefik
Sector
Voice AI platform
Audience
Teams that answer phones
Shape
Built, shipped and run by us
Stack
Next.js, FastAPI, Twilio
Next case studyVoxalyticsCall intelligence for sales and support teams: transcripts, AI scorecards and coaching from every recorded call.

Something like this to build?

Tell us what runs today and where it hurts. An engineer reads it and replies.