Build your own AI stack, or use a white-label one?

Building your own voice stack is cheaper per minute and more expensive per month. The vendor fees are public and small. The demo site, client portal, branding and intake are the parts nobody quotes you on.

Lumina8 min read
Scattered thin line fragments on a near-black field converging on the left into a single clean rounded rectangle outline lit by a warm amber glow.

Building your own AI stack is cheaper per minute and more expensive per month. A voice agent you assemble yourself runs roughly $0.10 to $0.15 a minute in vendor fees, which is nothing. What that figure leaves out is the demo site, the client portal, the branded subdomain, the usage metering, the intake process, and the two months of evenings you spend wiring them together. If you sell custom AI systems to local service businesses and your company is one to ten people, the real question is not whether you can build the stack. It is whether building it is the part of your business that makes money.

Here is the math, and the parts nobody quotes you on.

What does it actually cost to build your own voice stack?

Split the problem into two layers. There is the layer that has a published per-minute price, and there is the layer that does not. Almost every build-versus-buy argument online only counts the first one.

On the priced layer, the numbers are public. Retell publishes its components separately: voice infrastructure at $0.055 a minute, text to speech from $0.015 a minute on its platform voices and $0.040 with ElevenLabs, telephony at $0.015 a minute, and the language model anywhere from $0.003 a minute on a nano model to $0.32 a minute on a fast frontier tier. A sensible middle configuration lands near $0.13 a minute. Vapi prices it differently, charging $0.05 a minute for hosting with model provider costs passed through at cost, or nothing at all if you bring your own API keys. Both meter concurrency too. Vapi includes 10 lines and charges $10 per line per month after that, and Retell charges $8 per concurrent call per month past its free allotment.

If you skip the orchestration platforms and go straight to the carrier, Twilio charges $0.0085 a minute for inbound local calls, $0.0140 a minute outbound, and $1.15 a month for a local number.

Those rates also move under you. The per-minute floor has dropped all year, and we wrote up what GPT-Live-1's full-duplex voice changes the day after it shipped. Building your business on top of one vendor's rate card means rewriting your economics every time somebody launches.

So a client whose agent handles 500 inbound minutes a month costs you somewhere between $50 and $70 in vendor fees. That is a rounding error against any retail price that clears your wholesale cost. Which is exactly why per-minute cost is the wrong thing to optimize first. You are not going to win or lose this business on three cents a minute.

The layer with no per-minute price

Here is what you still do not have after you have wired up a voice agent that answers the phone and books an appointment:

A demo site a prospect can call before they have paid you anything. A client portal they log into after they sign. A subdomain with your logo on it instead of a vendor's. Usage caps that pause gracefully instead of producing a $900 surprise invoice. An intake process that collects hours, services, booking links, and phone routing without forty emails. Vertical-specific prompt packs so the agent for a dental office does not sound like the agent for a roofing company. Billing. A way to unpublish a demo when a deal dies.

None of that has a price per minute. All of it is the actual product your client is buying. The voice agent is the engine; everything else is the car.

There is also a maintenance cost that only shows up in production. OpenAI's own realtime cost guide explains that the full conversation is resent to the model on every turn, so later turns cost more than early ones, and that editing conversation history busts the prompt cache up to the point of the change. That is the kind of detail that turns a clean per-minute estimate into a bill that is 40% higher than you modeled. You find it by running real calls, not by reading a pricing page.

So what is the honest comparison?

Not per minute against per minute. It is a product you build against a product you resell.

On our side, the platform is $49 a month and wholesale delivery runs $225 to $500 a month depending on tier, with no build fee anywhere in the structure. You set retail, and everything above your delivery cost is yours. We wrote the margin math out in detail in how much to charge for an AI buildout, and the full tier breakdown is on the pricing section of the site.

The number that matters in a build-versus-buy decision is not the monthly fee. It is time to first paid client. If you bill $150 an hour for consulting, eighty hours of platform building is twelve thousand dollars of work you did not invoice, and you still have zero demos in front of prospects at the end of it. Buying converts that into a fixed monthly cost and gives you back the eighty hours to sell.

When building your own is genuinely the better call

We are not going to pretend this decision goes one way for everyone.

Build your own if you are running real volume. At tens of thousands of minutes a month, wholesale per-client pricing stops being efficient and direct vendor relationships win on cost. Build your own if you have an unusual technical requirement, like deep two-way sync with a client's proprietary scheduling system or a compliance posture that needs its own infrastructure. Build your own if you are actually a software company and the platform is the asset you intend to sell one day, because renting someone else's stack builds no equity for you. And build your own if you have a developer sitting idle, because then the labor cost you are avoiding is not real.

If you already have a working stack and a person who maintains it, adding a layer on top is a downgrade, not an upgrade. Do not do it.

What we are not good for

A few limits worth stating plainly, because you will find them anyway and it is better you find them here.

The $49 plan's usage allowance is a sales budget, not client production capacity. It covers 100 voice minutes and 500 chat messages a month, with a five minute cap per call and one published demo at a time. That is sized for demoing to prospects. Client delivery is the wholesale tier, billed separately, and if you go in expecting the $49 to also run your clients' live phone lines you will be annoyed on day three.

Sandbox disclosure always shows on a demo. The agent tells the prospect it is a demo and that bookings are simulated. We will not turn that off, including for you, because a prospect who discovers it themselves stops trusting the entire call.

You do not own the platform. If your business plan is to build software equity, this is the wrong tool and you should go build.

And a packaged vertical is a starting point, not a bespoke system. It gives you an agent that already speaks the trade's language on day one. It does not give you a custom integration into one client's twenty-year-old field service software.

How the decision usually goes

The real alternative to buying is usually not building. It is stalling. If you have a prospect who asked for a demo three weeks ago and a stack that is still 70% finished, you already know which problem you have.

The market timing argument is real here. The Census Bureau's Business Trends and Outlook Survey found national AI use among businesses hovering between 17% and 20% from December 2025 through May 2026, with under 20% of firms that have four or fewer employees reporting any AI use at all, against 37% of firms with 250 or more employees. Your buyers, the small local operators, are the least penetrated segment in the economy. That gap closes. It is a better use of the next quarter to be in front of those businesses with a working demo than to be in a text editor improving your own orchestration layer.

FAQ

Is it cheaper to build my own AI voice stack? Per minute, yes, and by a small margin that will not change your business. A middle-tier configuration on a platform like Retell or Vapi runs roughly $0.10 to $0.15 a minute all in. The cost that decides the question is the unpriced work: the demo, the portal, the branding, the intake, and your hours.

Can I use my own voice vendor with Lumina? Delivery runs on our infrastructure, so voice and chat are included with no API keys to manage and no phone numbers to buy. That is a feature if you want to sell rather than operate, and a limitation if you want vendor control. See the FAQ for the full detail.

What happens to my builds if I cancel? Published demos unpublish and everything you built stays as a draft. You keep signed-in access at zero usage, and resubscribing resumes it unchanged. Nothing gets deleted.

How long until I have something I can show a prospect? Around ten minutes for the first live demo, most of which is picking a subdomain and uploading a logo. You can try it against a real business name on the live demo before you decide anything.

Do I need to know how prompts work? No. The vertical packs are written. You will do better work if you understand them, and you can edit them, but the sale does not wait on your prompt engineering.

Where to start

Do not start with a spreadsheet. Pick one prospect you already know, in a trade you already sell to, and build a demo with their business name in it. Call it yourself. If that call is good enough to send to them, you have your answer about whether you need to build anything. If it is not, you have learned that in ten minutes instead of ten weekends.

Start a build, or read what happens after your client signs if you want the fulfillment side first.