Build voice agents that answer, call out, and hold a real conversation. One flat monthly price, unlimited minutes — you bring your own model, speech and telephony keys.
Every stage streams into the next — the agent starts speaking before the model finishes thinking. Interrupt at any moment and the pipeline flushes within one frame.
Bring your own providers & models
Configure the brain and the voice, place or receive real phone calls, and watch every conversation land in one dashboard.
Vocra knows when to listen, when to pause, and when to jump in. Interruptions cut the agent off mid-sentence, the way a real conversation flows.
Around 360ms voice-to-voice, measured in production. The reply starts streaming before the model has finished generating the rest of it.
Point an agent at a phone number and it dials out, handles the conversation, and hands you a transcript. Connect your own Twilio account.
Groq, OpenAI, Anthropic, Gemini, DeepSeek, Grok, Kimi or Qwen power the reasoning; Deepgram or ElevenLabs the voice. Swap providers per agent — your keys, your cost.
Every call recorded and transcribed with per-turn latency, so you can see exactly where the time went — not just that a call happened.
Unlimited minutes on every plan. You pay your model, speech and telephony providers directly at cost — Vocra bills only the platform.
A REST endpoint places the call; events and transcript come back over the API once it completes. SDKs are on the roadmap — today it's a plain HTTP call with a bearer key.
Read the docs →