GPT-Live-1 is out. What full-duplex voice changes for agencies
OpenAI put GPT-Live-1 into the API on September 10, 2026. It listens and speaks at the same time, and the voice layer costs $0.05 a minute. Here is what that changes for an agency selling AI phone systems to local businesses, and the four things it does not fix.

OpenAI put GPT-Live-1 into the API on September 10, 2026. It is a full-duplex voice model, which means it listens and talks at the same time instead of waiting its turn, and the voice layer costs $0.05 per minute, billed per second, with whatever backend model you point it at billed on top. If you sell AI phone systems to local service businesses, the short version is that the dead air after your prospect stops speaking is on its way out, and so is the objection you hear most often on demo calls.
That is the useful part. The rest of this is what actually changed, what it does to your cost per client, what it does not fix, and what I would say differently on a call this week.
What shipped on September 10
OpenAI's announcement describes GPT-Live-1 as a model that reasons over incoming and outgoing audio together, "avoiding the latency and brittle handoffs of chained STT-LLM-TTS architectures." The model had already been running inside ChatGPT. This release puts it behind an API endpoint.
Chained is how most voice agents have been built. Speech to text, then a language model, then text to speech, three services in a row, each adding delay. The caller finishes a sentence, everything goes quiet, and a beat later the agent starts talking. Every agency owner who has demoed a voice agent knows that beat. It is the moment the prospect's face changes.
Two numbers from the announcement are worth writing down. GPT-Live-1 "improves Full Duplex Bench performance by 30 percentage points over GPT-Realtime-2.1," and paired with GPT-6 Astra at medium reasoning effort it "ranks #1 on Tau3," which OpenAI describes as measuring voice-agent intelligence end to end. Those are OpenAI's own evaluations of OpenAI's own model, so weigh them accordingly. The more interesting figure comes from a customer: the language app Speak reported cutting interruptions by almost 80% against their previous turn-based system.
The model ships with twelve voices, native transcripts, keyword biasing, and turn detection that works without the system actually being turn-based. Tone and pace are steerable from the system prompt.
Why full duplex matters more for trades than for tech demos
A polished voice agent handles a clean, well-mannered caller. Real callers are not that.
A homeowner calling a roofing contractor after a storm talks over the greeting. Somebody phoning a plumber at 11pm says "hello? hello?" while the agent is still saying its opening line. People interrupt themselves, restart sentences, say "uh huh" in the middle of your agent's answer. A turn-based system treats all of that as a new turn and either barrels on or resets. A full-duplex model can, per the announcement, "respond to interruptions and acknowledgements as they happen."
That is the difference between an agent that survives one scripted demo question and an agent that survives a Tuesday. It is also the difference between a client who renews at month four and a client who quietly stops forwarding calls to it.
If you have never watched a prospect try to interrupt a voice agent, go do it on our live demo and interrupt it on purpose. That is the test buyers run without telling you they are running it.
What it costs, and what it does to your margin
Pricing moved to a shape that is much easier to quote.
The previous generation, gpt-realtime, went generally available in August 2025 at $32 per 1M audio input tokens and $64 per 1M audio output tokens. Audio tokens are close to impossible to estimate in advance, which is why so many agency owners have priced voice work by guessing and padding.
GPT-Live-1 is $0.05 a minute for the voice layer, billed per second, not rounded up. Backend model calls and tools bill separately at normal rates. So a client whose phone line runs 500 minutes a month costs you $25 in voice, plus the backend model, plus telephony. For comparison, ElevenLabs lists $0.080 per included call minute on its Agents plans, with LLM costs "based on usage" and telephony "at cost" on top.
None of those are retail prices. They are the floor under yours. The number your client sees has to cover the build, the account management, the month somebody changes their hours and forgets to tell you, and the fact that you are the one who picks up when it goes wrong. We went through that arithmetic in how much to charge for an AI buildout, and cheaper inference does not change the answer. It widens the gap between your cost and your price, which is the entire business.
Notice the pattern. Voice inference has gotten cheaper and better twice in about twelve months. If your pitch rests on having access to good voice models, your pitch has a short shelf life. Access is not scarce. Somebody who will configure it for a specific business, sit through the intake, and answer the phone in March still is.
What this does not fix
Four things, and I would rather you hear them from me than find them in week three.
Tool calls are not free and not instant. GPT-Live-1 delegates reasoning and tool calls to a backend text model. That is how it stays fast at the voice layer, but it means every calendar lookup and every CRM write goes through a second model that bills at its own rate and takes its own time. Your "book me Thursday at 2" round trip is only as quick as that backend.
Concurrency is capped by account tier. The model page lists concurrent Live sessions by usage tier, starting at 25 for Tier 1 and rising to 500 at Tier 5, with no free-tier access. That is fine for a handful of small clients. If you sell one business that runs a real call center, read that table before you sign.
The model does not know anything about your client. Knowledge cutoff is July 2025, image and video input are not supported, and nothing in the weights knows that this plumber does not do septic or that the med spa stopped offering a treatment in June. All of that comes from configuration, which comes from intake, which comes from a human getting answers out of a business owner who is on a roof. That handoff is still the slowest part of any build, and we wrote up the whole sequence in what happens after your client signs.
No local business is going to buy because of a model release. An HVAC owner does not know what full duplex means and will never search for it. They know they missed four calls on Saturday. The model release changes what you can build, not what makes them sign. If you find yourself explaining architecture on a sales call, you have already lost the thread.
What to change in your pitch this week
Three concrete edits.
Stop apologizing for the pause. If you have a line in your demo script that pre-empts the lag, cut it. Let the prospect run into a system that handles interruptions instead.
Hand them the phone earlier. The old sequencing, where you drove the demo and then let them try it, existed partly because turn-based agents fell apart when a stranger talked over them. That risk is smaller now. Let them call it in the first five minutes, in their own trade's language, with their own awkward question.
Change the proof you ask for. Instead of "what do you think," say "interrupt it, then ask it something it should refuse to answer." A buyer who watches a system handle being cut off and then decline to quote a price it does not have is a buyer who stops asking whether it sounds real.
Should you rebuild your stack for this?
If you already run your own voice infrastructure, understand session management, and enjoy that work, this is a migration worth evaluating on its merits. The economics are better than what you are running on.
If you do not, the honest answer is that a model release is a bad reason to start. The work that stands between you and a paying client is not the audio pipeline. It is 100+ vertical packs worth of trade-specific language, a demo your prospect can open without installing anything, a portal with your name on it, and a fulfillment path that does not require you to become an infrastructure company. That is the part we sell at $49 a month, with the voice and chat agents running on our side, no API key for you to obtain and no phone number for you to buy.
One limit on our own side. The $49 plan includes 100 live voice minutes a month, and single calls cap at five minutes. That is roughly forty demo calls. It is a demo allowance, not production call volume for a client, and anyone telling you a $49 subscription covers a client's live phone line is selling you something.
FAQ
Is GPT-Live-1 available to use right now? Yes. OpenAI's announcement says it is available in the API as of September 10, 2026, under the model string gpt-live-1. Custom voices require contacting their sales team. Free tier accounts do not get access.
Does this make gpt-realtime obsolete? Not officially. OpenAI's announcement uses GPT-Realtime-2.1 as a benchmark baseline and says nothing about deprecating it. The older model still has documented SIP telephony and MCP support, neither of which appears in the GPT-Live-1 announcement or its model page.
Will my clients notice the difference? Your clients will not. Their callers might. The gain shows up in messy conversations: people talking over the greeting, backchanneling, correcting themselves. If the only calls a business gets are short and orderly, the improvement is small.
Do I need to change anything in my Lumina demos? No. The voice and chat agents run on our infrastructure and the minutes are included in the subscription, so model changes on our side do not require anything from you. Your demo keeps working.
Is cheaper voice going to kill the margin on AI buildouts? It squeezes anyone whose only value was buying inference in bulk. It does not touch anyone selling configuration, delivery, and accountability, because those costs are human and they are not falling.
The next step
Open the live demo, pick the trade you sell into most, and do the thing your prospects do: call it and talk over the greeting. Then write down the two questions your last three prospects asked that the agent handled badly. Those two questions are your next demo script, and they matter more to your close rate than any model release.
If you want the demo on your own subdomain with your logo on it, that part takes about ten minutes: start here.
