What does a voice AI agent cost per minute?

A voice AI agent costs roughly $0.08 to $0.18 a minute in raw usage on the major platforms, and speech-to-speech models cost several times more. We did the math for a typical local service client and found the minutes are the cheap part.

Lumina8 min read
A thin amber waveform glowing across a dark charcoal background, peaking in a soft point of light above a row of fine meter tick marks

A production voice AI agent costs somewhere between about $0.08 and $0.18 a minute to run on the major platforms right now, before your time. That is the raw usage: the platform fee, the language model, the voice, and the phone line. Speech-to-speech models built for the most natural conversation cost several times more, and the minutes are almost never the expensive part of serving a client. Your hours are.

We price and sell voice agents every day, so we went through the current public price pages for Retell, Vapi, ElevenLabs and OpenAI this week and did the math for a typical local service client. Here is what we found, and what it means if you run a small AI agency.

What goes into the per-minute price of a voice agent?

Every voice call runs through four meters, whether the platform shows them to you or bundles them.

  1. Orchestration. The platform that handles turn-taking, interruptions, function calls and the call itself.
  2. The model. The LLM that decides what to say. This is the line that moves the most.
  3. The voice. Text-to-speech, unless you use a speech-to-speech model that handles listening and talking in one pass.
  4. Telephony. The phone number and the carrier minutes.

Retell publishes each of these separately, which makes it the easiest place to see the stack. Its pricing page lists voice infrastructure at $0.055 a minute, text-to-speech at $0.015 a minute for most voices and $0.040 for ElevenLabs voices, and telephony at $0.015 a minute through its Twilio or Telnyx connection. Phone numbers are $2 a month. The model sits on top, and Retell's headline range for a whole voice agent is $0.07 to $0.31 a minute.

What does a real client cost at those rates?

Take a plumbing or HVAC company that sends 800 minutes of calls a month to its AI receptionist. That is roughly 25 to 30 calls a day at a minute or so each, which is a busy small shop.

Using Retell's published components:

SetupPer minute800 minutes
Lean: standard voice, GPT 5.4 mini ($0.024)$0.109$87.20
Premium: ElevenLabs voice, Claude 5 Sonnet ($0.064)$0.174$139.20
Speech-to-speech: GPT Realtime 2.1 model line alone ($0.38)$0.38+$304+

The first two rows add infrastructure, voice, model and telephony. The third row is only the model line Retell lists for GPT Realtime 2.1, before infrastructure and phone minutes, so the real figure lands higher.

Vapi comes out in the same neighborhood by a different route. Its pricing page charges $0.05 a minute for hosting and passes model costs through at cost. Its own calculator puts 1,000 minutes at $82 to $129 a month, of which $50 is the hosting fee. That works out to $0.08 to $0.13 a minute.

ElevenLabs bundles differently. Its Agents pricing includes a block of minutes with each plan, 1,238 minutes on the $99 Pro plan, and charges $0.08 a minute for calls beyond that, rising to $0.16 at burst. The LLM and telephony are billed on top.

So for most small-business receptionist work, the raw usage bill for a single client is under $150 a month. That is the number your prospects' nephews quote when they say "AI costs pennies." They are right about the minutes. They are wrong about everything else.

Why do speech-to-speech models cost so much more?

Speech-to-speech models skip the separate transcription and voice steps and work on audio directly. The payoff is better timing, fewer awkward pauses, and more natural interruptions. The cost is steep.

OpenAI's API pricing page lists GPT-Realtime-2.1 at $32 per million audio input tokens and $64 per million audio output tokens. The mini version is $10 and $20. Retell's per-minute translation of those is $0.38 for the full model and $0.07 for the mini.

That gap matters. The full model is five times the price of the mini for the same minute. For a dental office booking cleanings, we don't think the full model earns its cost. For a law firm intake line where a caller is upset and talking over the agent, it might.

Google made this more interesting last week. The Gemini API changelog shows Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking reaching general availability on September 15. Google describes the first as the default for low-latency voice agents, with asynchronous function calling on by default. The changelog entry doesn't give a price, so we won't guess one. The point for an agency is simple. There are now two major vendors shipping production audio-to-audio models, and that tends to pull prices down over the following months. We covered the OpenAI side of this earlier in our post on GPT-Live-1 and full-duplex voice.

If minutes are cheap, where does the money actually go?

This is the part that most per-minute articles skip, because most of them are written by platforms selling minutes.

Getting a voice agent to production for one client means writing and testing the prompt for their trade, loading their services and price bands, wiring the booking tool to their calendar, setting up call transfer and text-back, and then listening to calls for the first two weeks and fixing what breaks. After launch, somebody has to update hours for holidays, add a new service, and answer the owner when he says "it told a customer we do gas lines, we don't."

None of that shows up on a price page. On a small account, it dwarfs the $90 in minutes.

It also explains why the pricing conversation with your client should never be anchored on usage. If you quote "$0.15 a minute plus a margin," you have invited the client to compare you against a platform they could sign up for themselves. Price the outcome and the ongoing work instead. We laid out how in how much to charge for an AI buildout.

Where the per-minute math breaks down

A few honest limits on the numbers above.

Token-priced models make per-minute figures estimates. OpenAI prices its realtime models by tokens, and a chatty caller or a long system prompt burns more tokens per minute. Retell's per-minute conversion is a reasonable average, not a guarantee.

Fixed fees add up at scale. Retell charges $8 per concurrent call slot per month after the first 20. Vapi's Pro tier has a $999 monthly minimum. ElevenLabs caps concurrency by plan. If you manage 30 clients on one account, check concurrency before you check minutes.

Add-ons are real. Retell lists knowledge base, denoising and safety guardrails at $0.005 a minute each, PII removal at $0.01, and SMS at $0.01 a message. Turn on everything and the lean setup gets less lean.

Prices change often. Every figure here came from the vendors' own pages this week. Recheck before you build a pricing model on them.

And a note on who should care. If you sell a handful of clients and use a white-label stack, the per-minute cost is mostly someone else's problem, and you should spend your time selling. If you are assembling your own stack across Retell or Vapi, a model provider and a phone carrier, these numbers are your margin, and you should know them cold. We compared both routes in build your own AI stack or use a white-label one.

How we handle usage at Lumina

On our side, the voice and chat agents run on our infrastructure. There is no API key to get, no phone number to buy, and no metered invoice on any plan. The $49 plan covers your demo site and includes 100 live voice minutes, 500 chat messages and 50 website audits a month for demos, which is about forty demo calls. When you close a client, delivery goes through the fulfillment platform at a flat wholesale monthly. The top tier, which includes Voice AI, is $500 a month, and you set the retail price. The full list of limits sits on the pricing section.

The tradeoff is control. You don't pick the model per client or tune per-minute costs yourself. For most agencies with under ten clients, that is the right trade. If you want to squeeze every cent out of every minute, build your own stack.

FAQ

How much does an AI receptionist cost per month in usage?

For a small service business sending 800 minutes a month through Retell, the raw usage lands between about $87 and $140 depending on the model and voice. That excludes setup, monitoring and changes, which usually cost more than the minutes.

Is Retell or Vapi cheaper?

They land close together. Retell's itemized components put a lean agent near $0.11 a minute. Vapi's own calculator puts 1,000 minutes at $82 to $129, or $0.08 to $0.13 a minute. Your model and voice choice matter more than the platform.

Should I use a speech-to-speech model for my clients?

Only where conversation quality directly affects revenue, such as legal intake or high-ticket consultations. The full GPT Realtime 2.1 model line costs about five times the mini version per minute on Retell. For routine booking calls, a standard pipeline is fine.

Should I bill clients per minute?

We don't recommend it. Per-minute billing invites the client to compare you against the platform's price page and ignores the work that makes the agent good. Charge a flat monthly for the outcome and build usage into it.

What to do next

Pull up your last three client quotes and write the raw usage cost next to each one using the table above. If usage is more than 15 percent of what you charge, your pricing is anchored in the wrong place. If you would rather not track minutes at all, try the live demo and then start a workspace for $49 a month. Prices go up October 1.