That's about 700 three-minute calls. Set that next to a part-time receptionist and next to the bookings that currently die in voicemail. The rest of this article is where every cent of the per-minute cost goes, what moves the build fee, and when Vapi or Retell is the cheaper choice. KeenCraft has shipped on Vapi, Retell, and LiveKit. This is the math we walk through on a strategy call.
Where each minute of a call actually goes
Every AI voice call is five services billing at once. Platforms advertise one of them as the headline price, which is why the number on the pricing page rarely matches the invoice. The rates below are US list prices checked in September 2026. Providers change them, so confirm the live page before you quote a client.
| Layer | What it does | Typical provider | Cost per minute |
|---|---|---|---|
| Telephony | Connects the agent to a real phone number | Twilio | ~$0.0085 inbound, ~$0.014 outbound |
| Speech-to-text | Turns the caller's voice into text in real time | Deepgram Nova-3 | ~$0.0048 streaming (promotional; regular list $0.0077) |
| LLM | Decides what to say and which tool to call | OpenAI, Anthropic, Google | ~$0.003 to $0.16, by model |
| Text-to-speech | Speaks the reply | Cartesia, ElevenLabs | ~$0.015 standard, ~$0.04 ElevenLabs |
| Orchestration | Runs the conversation loop, turn-taking, and interruptions | Vapi, Retell, or LiveKit Cloud | $0.05 Vapi, $0.055 Retell, $0.01 LiveKit |
Add a standard managed stack and you land around $0.11 to $0.15 per minute: Twilio inbound, Deepgram, a fast mid-size model, a $0.015 voice, and Vapi or Retell on top. The same components on LiveKit Cloud come in near $0.07 per minute, because orchestration drops from about five cents to one cent after the plan's included agent minutes. LiveKit bills $0.01 per agent-session minute once that allotment is used. Below the allotment, orchestration is cheaper still.
Two layers decide most of the bill: the voice and the model. A premium voice can cost nearly three times a standard one. A frontier model can cost many times a small, fast one. For appointment booking, a fast mid-size model is usually the right call. Callers notice a slow reply long before they notice a slightly less clever sentence. The production bar for that reply, including the calendar write, is in how to build an AI voice agent that actually books appointments.
What pushes the build fee up or down
The per-minute cost is mostly set by the vendors above. The build fee is where projects differ, and it follows how much of the business the agent has to touch.
| Factor | Lower end | Higher end |
|---|---|---|
| Call direction | Inbound only: answer, qualify, book | Inbound plus outbound follow-up, with retry rules and calling-hour limits |
| Calendar and CRM | Books into one calendar | Reads availability and writes contacts, stages, and notes to HubSpot, Salesforce, or Zoho |
| Locations | One business, one phone number | Multiple locations, each with its own hours, staff, and isolated data |
| Languages | English only | Several languages, or one with thin speech-model support |
| Compliance | General business calls | Healthcare calls that need HIPAA-conscious data handling |
| Handoff | Takes a message when stuck | Warm transfer to a live person with the context attached |
At KeenCraft, a single-location inbound agent that books into one calendar usually costs $3,000 to $12,000. A multi-location agent with CRM write-back and outbound follow-up usually costs $12,000 to $30,000. Either way, the total is in a written Statement of Work before any code is written, and each milestone is tested before it is paid.
| Price | What you get |
|---|---|
| ~$3,000 | One location, inbound only, English, books into Google Calendar or Calendly, takes a message when it cannot help |
| ~$12,000 | Still one location, connected to the practice management or booking system, HIPAA-conscious data handling, warm transfer to staff, or a second language |
| $12,000 to $20,000 | Two to five locations, each with its own hours and staff, plus contact and stage write-back to HubSpot or Zoho |
| ~$30,000 | Many locations with isolated data per location, Salesforce custom objects, and outbound follow-up with retry rules and calling-hour limits |
Those ranges are wide because the rows above are real scope, not padding. After one 30-minute call you get a single fixed number in writing, tied to the items in this table. If the work sits at the $3,000 end, that is the price. If a platform subscription would serve you better than a build, we say that on the call.
The biggest hidden driver is usually the CRM, not the voice. A call that sounds good is the smaller piece of engineering. Creating the right contact, skipping duplicates, and moving the right pipeline stage is where the hours go. That write-path is CRM automation for HubSpot, Salesforce, and Zoho. Outbound follow-up, with its retry rules and calling-hour limits, is the other step up, covered in outbound follow-up after a missed call.
Multi-location work costs more because each site needs its own hours, staff, and data boundary. Dynaris is the live version of that pattern: tools isolated per workspace, so one location cannot book another location's calendar. How routing, calendars, and that boundary work is in AI front desk for multi-location businesses. Voice, chat, and email on one thread is an AI front desk that is more than a phone agent. Healthcare data handling is the same kind of scope. OptimateMD.health was built with those standards in the architecture. On Vapi, HIPAA is a published add-on at $2,000 a month on top of usage, which belongs in the run-cost comparison when the calls are clinical.
When Vapi or Retell is cheaper than a custom build
Use Vapi or Retell when you handle under a few thousand minutes a month, you are still testing whether callers will stay on the line with an AI agent, or you have no one to run your own voice infrastructure. Both can be live in days. At low volume their platform fee is a small price for skipping the engineering. No-code, enterprise, hybrid, and a studio build are the other choices, mapped in how to choose an AI appointment booking agent.
A custom build on LiveKit starts to pay off as volume grows. The orchestration fee drops from about $0.055 on Retell to about $0.01 on LiveKit Cloud, roughly $0.045 saved on every minute once you are past the included allotment:
| Monthly call minutes | Retell at $0.055 | LiveKit Cloud at $0.01 | Monthly saving |
|---|---|---|---|
| 2,000 | $110 | $20 | $90 |
| 10,000 | $550 | $100 | $450 |
| 50,000 | $2,750 | $500 | $2,250 |
Divide the extra cost of the custom build by that monthly saving and you have a payback period. At 10,000 minutes a month, the orchestration savings alone can cover a smaller build inside the first year, before the other reasons to own the stack. A $12,000 build takes longer than that on orchestration savings alone. At 2,000 minutes, a Retell or Vapi subscription is usually the honest recommendation. Agencies rebilling that usage across many clients have a different cutoff, covered in white-label AI voice agents for agencies.
Latency is part of what you are buying
The other reasons are often why a business goes custom anyway. You control latency. On Rawk.ai, a voice agent builder KeenCraft leads development on, end-to-end response time dropped from about 900ms to about 320ms by owning the pipeline. You also keep the call data, choose each provider, and avoid rebuilding the agent if a platform changes its pricing. VoiceCake is the appointment proof on the other side of that choice: a live inbound platform for dental, healthcare, fitness, and mortgage clients, including an after-hours receptionist that books onto the real schedule.
Costs that show up on the first invoice
- Testing minutes. Every test call bills like a real one. Tuning an agent before launch can burn hundreds of minutes.
- Calls that go nowhere. Voicemails, hang-ups, and spam still use telephony, speech-to-text, and orchestration.
- Transfers. A warm transfer keeps two call legs open, so telephony roughly doubles for those minutes.
- Monitoring. Recording, transcripts, and observability are often billed separately, commonly around $0.01 per minute.
- Phone numbers and concurrency. Each number has a monthly fee, and a busy hour can hit a platform cap on simultaneous calls. Vapi includes 10 concurrent lines, then about $10 per extra line each month. Retell includes 20.
- Ongoing tuning. Prompts need adjusting once real callers say things nobody scripted. Agree up front who owns that work and what it costs.
None of these is large alone. Together they are why a realistic budget adds 15 to 20 percent on top of the raw per-minute math for the first two months. The buffer usually shrinks after launch, once testing stops and the prompts settle.
Get a number for your call volume
The fastest way to price an agent is to run your own minutes through this math. Book a free 30-minute strategy call and KeenCraft will estimate the per-minute cost, say whether Vapi or Retell is enough for now, and give a fixed build price when custom is the better buy. A written proposal typically lands in 2 to 5 business days.
Sources
Prices checked September 2026. Vendor pages move, so verify before you put a rate in a proposal.
- Retell AI pricing: voice infrastructure $0.055/min, most voices $0.015/min, ElevenLabs $0.040/min, advertised all-in range about $0.07 to $0.31/min.
- Vapi pricing: hosting $0.05/min. Speech-to-text, the model, the voice, and telephony are billed by those providers, or at no Vapi charge if you bring your own keys.
- LiveKit pricing: agent session minutes at $0.01/min after the plan allotment, and Deepgram Nova-3 monolingual at $0.0048/min on the Build and Ship plans.
- Deepgram pricing: Nova-3 monolingual streaming, current promotional rate $0.0048/min, regular price $0.0077/min.
- Softcery voice agent cost calculator: Twilio US local about $0.0085/min inbound and $0.014/min outbound, checked alongside the platform rates above.