Last updated: 1 October 2026
Quick Answer: Retell AI pricing is pay-as-you-go per connected minute, summed from voice infrastructure, text-to-speech, LLM and telephony charges. Concurrency beyond the 20 included lines, extra phone numbers and knowledge bases are billed monthly on top. Enterprise agreements replace the public rate with a contracted one and lift the concurrency cap.
The published rate is a range, not a price, and the gap between its ends decides whether a pilot survives budget review.
How much does Retell AI cost in 2026?
Retell AI's own documentation puts a production minute at $0.07 to $0.31, read on 1 October 2026. Voice infrastructure, text-to-speech and the LLM sit inside that figure; telephony is added on top.
Every minute is tracked to the nearest second, and where you land in that band is set by the model and the voice, not a plan tier. Retell AI competitors meter different bundles, so compare on a worked month.
| Line item | What drives it | How to estimate yours |
|---|---|---|
| Conversation voice engine | Retell AI's infrastructure, every connected minute | Minutes times the engine rate |
| Text to speech | The voice provider you pick | A premium voice costs nearly three times a platform one |
| LLM | The model, and whether the fast tier is on | Widest swing on the invoice |
| Telephony | A Retell AI number, or your own SIP trunk | Your own trunk moves it off this bill |
| Concurrency, numbers, knowledge bases | Peak simultaneous calls, caller ID | Size to peak, never average |
| Billing adjustments | Dynamic openings, long prompts | Audit prompt length first |

How is Retell AI's per-minute price built up?
On the rate table read on 1 October 2026, Retell AI's pricing page lists the conversation voice engine at $0.055 a minute. Platform, Cartesia, OpenAI and Inworld voices cost $0.015 and ElevenLabs voices $0.040. Telephony is $0.015 on a Retell AI number and nothing on your own SIP trunk. Models run from $0.0016 a minute at the small end to $0.64 for the priciest fast tier, and GPT 5.6 Terra and Claude 5 Sonnet, two of four recommended models, cost $0.064. The monthly items are $2.00 per phone number, $10.00 per verified number (one table calls it a one-time fee, so confirm it), $8.00 per concurrent line beyond the first twenty and $8.00 per knowledge base beyond the first ten.
What drives Retell AI cost per minute
Add the four usage components and you have your rate. Two rules add billed seconds beyond the call's clock time: a call under ten seconds that opens with a dynamic message is billed ten seconds, and a prompt past 4,000 tokens multiplies billed duration by its token count divided by 4,000.
What do concurrency, phone numbers and knowledge bases add?
These are the lines teams forget, because they hit a card at the end of the cycle instead of drawing down usage credits. A pay-as-you-go workspace starts with twenty lines, and inbound calls above the quota wait about 40 seconds, then go to a fallback number if one is set.
Phone numbers are billed monthly per number; verified caller ID and branded calling are priced separately, because carriers flag unverified outbound numbers as spam. Knowledge bases are free up to ten, then charged per base per month, and retrieval adds a small per-minute fee. Denoising, guardrails and PII removal each add a cent a minute or less.
What does a 10,000-minute month cost on Retell AI?
This is our own arithmetic, not a Retell AI quote. Take 10,000 connected minutes on a Retell AI number, two numbers and fifty concurrent lines. With the $0.055 engine and $0.015 telephony, a platform voice and GPT 5.4 mini works out near $0.11 a minute, a platform voice and GPT 5.6 Terra near $0.15, and an ElevenLabs voice and GPT 5.5 near $0.27; that is $1,090, $1,490 and $2,700 of usage, and the thirty extra concurrent lines plus two numbers add about $244 a month, so the same traffic costs roughly $1,330 to $2,950 depending on choices made in a dropdown. Burst capacity, if enabled, adds $0.10 a minute to every call that starts above quota.
To build your own figure:
- Forecast connected minutes, including transfer time on the line.
- Pick the voice and the model, and read both rates off the rate table.
- Add the engine and telephony rates, or zero telephony for your own trunk.
- Multiply by minutes, then add concurrency, numbers and caller ID.
- Add contingency for the ten-second minimum and prompt-token scaling.
What mistakes should you avoid when estimating Retell AI cost?
Four errors explain most of the gap between forecast and invoice. Sizing concurrency to average rather than peak drops inbound calls or triggers a burst surcharge. Pricing the demo model and shipping a different one moves the usage line by multiples. Silence is billed, so a forecast built on talk time understates hold-heavy queues. And a transfer is not free: the agent fee stops, telephony does not.
How does Retell AI pricing compare with Vapi and Bland AI?
The three meter different bundles, so headline rates are not comparable. Vapi charges $0.05 a minute for hosting and passes transcription, model and voice through at provider cost. On top sit optional support packages, $29 monthly at the low end and a tenth of the hosting fee with a $999 monthly minimum, plus $2,000 a month for HIPAA-eligible data handling and $10 per extra concurrent line. Its own calculator put 1,000 minutes at $82 to $129 on 1 October 2026. Bland AI, read 1 October 2026, quotes $0.14 a minute with no platform fee, or $0.12 a minute with a $299 monthly platform fee, and folds the model, speech-to-text and text-to-speech into that number while billing telephony separately. Searches for Retell AI vs ElevenLabs mix layers: ElevenLabs is a voice option inside Retell AI's stack and also sells a rival agent platform, ElevenAgents.
What are the alternatives to Retell AI?
Retell AI alternatives split three ways. Component-priced platforms such as Vapi expose every provider rate, which suits teams that already negotiate model and speech contracts. All-inclusive platforms such as Bland AI fold the model and speech layers into one rate, which simplifies forecasting but gives you less say over what sets the cost.
Enterprise-contract vendors such as Synthflow list only contracted pricing. Its pricing page, read 1 October 2026, scopes contracts around call volume, concurrency, telephony, integrations and launch support. Choose a metered platform when volume is uncertain, an all-inclusive rate when you want one predictable number, and an enterprise contract when security review outweighs speed. Our comparison of voice agent platforms covers the capability differences.
When is a self-hosted voice stack cheaper than Retell AI?
A self-serve metered plan has little fixed cost and a flat rate, though Retell AI's enterprise tier advertises volume pricing; a stack you run has fixed infrastructure and engineering cost and a marginal rate near zero. Take your blended rate, multiply by forecast minutes over a year, then compare with compute, a speech layer, carrier minutes and the time to run it. Predictable volume favors the stack; spiky volume favors the meter.
When a Retell AI voice agent moves in-house
Two conditions usually decide it before cost does. If call audio cannot leave your network, hosted Retell AI cannot meet that on settings alone; its homepage lists an enterprise on-prem option, so ask sales. And owning the stack lets you swap the speech provider or fine-tune a domain model.
How Origins AI approaches voice agent cost at scale
Origins AI (originshq.com) does not publish a rate card for its voice product, so there is no per-minute figure to set beside the table above. Origins AI Voice AI is deployed inside the customer's own environment: on-premise, in a private cloud on your own AWS, Azure or GCP account, in a hybrid mode, or air-gapped.
The page lists no metered rate, so ask any quote to separate infrastructure, implementation and support; concurrency becomes capacity you own, at up to 500 sessions per node, per the product page.
Speech providers and models are yours to bring, so the most volatile lines above move to contracts you already hold, and telephony connects over SIP trunking, PSTN or Twilio.
In on-premise and air-gapped modes, the page states that no calls are routed through a third-party cloud and no data leaves your network. Origins AI reports a pilot live in 30 days, with the same team building the surrounding integrations into CRM and ticketing.
Talk to an engineer
Share your call volume, peak concurrency and data constraints, and we will map metered cost against a deployment in your own network. Book a technical call.


