A text-to-speech API looks like the easy half of a voice product: send a string, get audio back. In production the bill is metered per character, every pause before the first word reads as lag, and the model decides on its own whether "Dr." means doctor or drive.
Three things moved in the past two weeks. On September 28, 2026, ElevenLabs shipped Eleven v4 at $0.08 per 1,000 characters. On October 1, OpenAI deprecated every model behind its speech endpoint, with a shutdown date of January 6, 2027. And on October 6, when we fed two popular open-source models the same appointment reminder, both read "Tue." as the number two.
How We Ranked and Priced Them
"Best" here means the Artificial Analysis Provider Voice Arena, a 92-model board where listeners pick the better of two clips without knowing which model made each, scored as Elo ratings as of October 6, 2026. A separate pronunciation test has reviewers check numbers, dates and codes.
Most APIs bill per character of input, so your own string is the meter; a few bill tokens, so their cost follows audio length. Every price below, checked on vendor pages on October 7, 2026, is for one workload: a US clinic's reminder line speaking 5 million characters a month, about 39,000 calls of this 128-character message:
"Your appointment with Dr. Patel is confirmed for Tue., Oct. 13 at 2:30 PM. Your copay is $45.50. Questions? Call (555) 010-1234.
What Two Open Models Did With One Appointment Reminder
We ran the reminder through Kokoro 0.9.4 and Piper 1.8.0 on an M3 Pro MacBook Pro's CPU, then transcribed each WAV file back with faster-whisper 1.2.1 and its small.en model.
| Model | Warm synthesis | Audio length | Real-time factor | What came back |
|---|---|---|---|---|
Kokoro-82M, af_heart voice | 1.52 s | 14.22 s | 0.107 | "confirmed for 2. Oct. 13", "Call 555-1012-34" |
Piper, en_US-lessac-medium voice | 0.347 s | 13.66 s | 0.025 | "confirmed for 2.00. Oct. 13", "Call 555-010-1234" |
Both models spoke "Tue." as the number two, and Kokoro regrouped the phone number into something a patient cannot dial. The fix was a dozen lines of Python that expand abbreviations and read phone numbers digit by digit before the text reaches any engine:
Pythonimport re DAYS = {"Mon": "Monday", "Tue": "Tuesday", "Tues": "Tuesday", "Wed": "Wednesday", "Thu": "Thursday", "Thurs": "Thursday", "Fri": "Friday", "Sat": "Saturday", "Sun": "Sunday"} MONTHS = {"Jan": "January", "Feb": "February", "Mar": "March", "Apr": "April", "Jun": "June", "Jul": "July", "Aug": "August", "Sep": "September", "Sept": "September", "Oct": "October", "Nov": "November", "Dec": "December"} def alternation(words): return "|".join(sorted(words, key=len, reverse=True)) def normalize_for_tts(text: str) -> str: text = re.sub(rf"\b({alternation(DAYS)})\.,?", lambda m: DAYS[m[1]] + ",", text) text = re.sub(rf"\b({alternation(MONTHS)})\.", lambda m: MONTHS[m[1]], text) text = re.sub(r"\bDr\.(?= [A-Z])", "Doctor", text) # US phone numbers: one digit at a time, with a pause between the groups text = re.sub(r"\(?\b(\d{3})\)?[ .-]?(\d{3})[ .-](\d{4})\b", lambda m: ", ".join(" ".join(group) for group in m.groups()), text) return text if __name__ == "__main__": line = ("Your appointment with Dr. Patel is confirmed for Tue., Oct. 13 at 2:30 PM. " "Your copay is $45.50. Questions? Call (555) 010-1234.") print(normalize_for_tts(line))
textYour appointment with Doctor Patel is confirmed for Tuesday, October 13 at 2:30 PM. Your copay is $45.50. Questions? Call 5 5 5, 0 1 0, 1 2 3 4.
With normalized input, both came back as "Tuesday, October 13" and "Call 555-010-1234", and hosted models need the same help: Eleven v4 scores 91.7 percent on the pronunciation test, Cartesia Sonic 3.6 74.5.
ElevenLabs Eleven v4 and v4 Turbo
ElevenLabs holds the top two spots: Eleven v4 Turbo at 1334 Elo and Eleven v4, launched September 28, 2026, at 1321.

Both cover 90-plus languages with voice cloning and lead the pronunciation test at 91.7 and 90.1 percent. v4 Turbo has a median inference latency of about 100 milliseconds, measured on ElevenLabs' servers rather than your network. The catch is price: Eleven v4 costs $400 for our workload and caps a request at 10,000 characters.
Pricing: Eleven v4 $0.08 per 1,000 characters · v4 Turbo and Flash v2.5 $0.04.
Alibaba Qwen-Audio TTS
Qwen-Audio-3.1-TTS-Plus from Alibaba ranks third at 1292 Elo, and Qwen-Audio-3.0-TTS-Plus ranks sixth at 1261.

Alibaba's docs say the Plus model improves low-resource languages and Chinese dialects. For US teams the limits are practical: Model Studio lists Beijing and Singapore deployments, with no US region, and the international pricing page publishes a rate for 3.0 but not yet for 3.1.
Pricing: Qwen-Audio-3.0-TTS-Plus $0.20 per 10,000 characters in Singapore · 3.0-TTS-Flash $0.15. Our workload: $100 on 3.0 Plus.
Cartesia Sonic-3.6
Cartesia ranks fourth at 1278 Elo with Sonic-3.6, a 44-language model it advertises at sub-90-millisecond latency.

Its best feature is versioning: sonic-3.6 tracks the latest stable snapshot, while sonic-3.6-YYYY-MM-DD pins one that never changes. It scored 74.5 percent on pronunciation despite claiming to read confirmation codes correctly, and plans cap concurrent requests at 2, 3, 5 and 15 from Free to Scale.
Pricing: Startup $49/month for 1.25 million · Scale $299 for 8 million · overage $38 to $65 per million. Our workload: $217.75 on Startup with overage.
Google, Amazon and Microsoft
The clouds are easy to buy on existing contracts, but their cheapest voices rank low.

Google Cloud Text-to-Speech has the fifth-ranked model, Gemini 3.8 Flash TTS at 1275 Elo and 89.5 percent on pronunciation, billed at $0.50 per million text tokens in and $9 per million audio tokens out through December 31, 2026, doubling on January 1, 2027. Chirp 3 HD voices, 58th, cost $30 per million characters with the first million free each month.
Amazon Polly charges $30 per million characters for Generative voices, 53rd, and $16 for Neural, 90th of 92. Azure AI Speech charges $22 for Neural HD, 28th, and $15 for Neural, 68th, in East US per Microsoft's Retail Prices API.
Inworld TTS-2
Inworld ranks seventh at 1251 Elo with Realtime TTS-2, and its cheaper TTS-2 Flash ranks twelfth at 1214.

Inworld serves OpenAI's POST /v1/audio/speech in OpenAI's own request format, so migrating means changing the SDK base URL to https://api.inworld.ai/v1 and the model to inworld-tts-2. Its docs list 200-plus languages.
Pricing: TTS-2 $25 per million characters on demand, $17.50 on Builder ($100/month) · TTS-2 Flash $15 on demand, $9 on Builder. Our workload: $125 on demand, $100 on Builder.
Speechify Simba 3.2
Speechify's developer API, sold separately from its reader app, ranks eighth with Simba 3.2 at 1242 Elo.

Its $10 per million characters on Starter, $8 on Pro and $6 on Scale is the lowest published per-character rate in the top ten, and Speechify cites independent first-audio times of 106 milliseconds on Coval and 123 on Voice Arena (September 24, 2026). The limit is language: self-serve accounts get English on Simba 3.2 and six languages on the older Simba 3.0.
Pricing: 500,000 characters free each month · Starter $10/month with 1.9 million included · Pro $99 with 13.5 million · Scale $499 with 78 million. Our workload: $41 on Starter.
MiniMax Speech 2.8
MiniMax ranks 19th with Speech 2.8 HD at 1173 Elo and 22nd with Speech 2.8 Turbo.

Its strengths are long-form audio and cloning: the asynchronous endpoint takes up to 1 million characters per request, and rapid voice cloning costs $1.50 per voice. It is also the most expensive API here per character.
Pricing: Speech 2.8 HD $100 per million characters · Speech 2.8 Turbo $60 · audio subscriptions from $5/month. Our workload: $500 on HD, $300 on Turbo.
Fish Audio S2.1 Pro
Fish Audio ranks 24th with S2.1 Pro at 1141 Elo.

The API has no subscription or monthly minimum, and concurrency rises from 5 to 15 to 50 requests as your total spend passes $100 and $1,000. It bills UTF-8 bytes rather than characters, so English text costs a per-character rate while accented or CJK text costs more.
Pricing: S2.1 Pro $15 per million UTF-8 bytes. Our workload: $75.
Deepgram, Hume, Rime and the Rest
Deepgram, absent from the board, launched Flux TTS on August 12, 2026: it keeps the whole call in context, reports what the caller heard when they interrupt, quotes time to first audio as low as 80 milliseconds and can run self-hosted. Aura-2 costs $0.030 per 1,000 characters and Flux TTS $0.045.

Hume Octave 2 ranks 61st, but Hume's pricing page now redirects to its homepage (checked October 7, 2026), so ask its sales team for a rate. Rime, whose Coda model ranks 55th, charges $0.05 per 1,000 characters for Coda and $0.03 for Mist v3, and sells on-premises deployments.
Luna TTS from VUI Labs (10th), Soniox (17th), Smallest.ai Lightning V3.1 Pro (20th) and Murf Falcon 2 (21st) also made the top 25, but we could not verify a per-character price on their pages; Soniox estimates about $0.70 per hour of generated speech.
OpenAI's Speech Endpoint Shuts Down January 6, 2027
OpenAI's deprecations page lists tts-1, tts-1-hd and both gpt-4o-mini-tts snapshots for shutdown on January 6, 2027, in a notice dated October 1, 2026, and names gpt-realtime-2.1-mini as the replacement.

That replacement runs on the Realtime API, a WebSocket or WebRTC session rather than a REST call, at $20 per million audio output tokens. Until the shutdown, tts-1, 41st, costs $15 per million characters. Do not start a new build on /v1/audio/speech; if you have one, Inworld's compatible endpoint is the smallest code change.
Open Models: Kokoro, Piper, Chatterbox and Breeze
Open weights trade the meter for setup work, and the license decides what you can ship.

Kokoro-82M ranks 54th and is Apache 2.0 licensed. Piper ran 40 times faster than real time on our CPU, but its MIT repository was archived on October 6, 2025, and development continues as piper1-gpl under GPL-3.0. Chatterbox from Resemble AI is MIT licensed, and its Multilingual V3 clones voices in 23-plus languages. Orpheus from Canopy Labs is Apache 2.0.

The best-ranked open weights, BreezeBlue's Breeze TTS 2 at 11th and Fish Audio's S2 Pro, are licensed for research and non-commercial use only. Kokoro and Piper also bundle espeak-ng, itself GPL-3.0; on macOS we had to point ESPEAK_DATA_PATH at a short path to its espeak-ng-data folder before either would run.
Side by Side
| API or model | Arena rank | Per 1M characters | 5M characters/month | Self-host |
|---|---|---|---|---|
| ElevenLabs v4 Turbo | 1 | $40 | $200 | No |
| ElevenLabs Eleven v4 | 2 | $80 | $400 | No |
| Cartesia Sonic-3.6 | 4 | $45 past Startup's 1.25M | $217.75 | No |
| Google Gemini 3.8 Flash TTS | 5 | Token-billed | Token-billed | No |
| Qwen-Audio-3.0-TTS-Plus | 6 | $20 | $100 | No |
| Inworld TTS-2 | 7 | $25 | $125 | No |
| Speechify Simba 3.2 | 8 | $6 to $10 | $41 | No |
| MiniMax Speech 2.8 HD | 19 | $100 | $500 | No |
| Fish Audio S2.1 Pro | 24 | $15 per M bytes | $75 | No |
| Azure AI Speech Neural HD | 28 | $22 | $110 | No |
| OpenAI tts-1 | 41 | $15 until Jan 6, 2027 | $75 | No |
| Google Chirp 3 HD | 58 | $30, first 1M free | $120 | No |
| Amazon Polly Neural | 90 | $16 | $80 | No |
| Deepgram Aura-2 | Not ranked | $30 | $150 | No |
| Kokoro-82M | 54 | $0 | Hardware only | Apache 2.0 |
Ranks are Artificial Analysis Voice Arena positions out of 92 models on October 6, 2026; prices are list rates checked October 7, 2026.
How to Choose Without Migrating Twice
- Count your real characters. Export a month of the text you synthesize, SSML included, and multiply by the rates above.
- Normalize, then compare. Run your 50 hardest strings, such as order numbers and drug names, through
normalize_for_ttsand each candidate, then round-trip the audio through speech-to-text. - Time the first audio byte from your region,
us-east-1for most US stacks, because vendor latency is measured on their side. - Keep an exit. Pin dated model snapshots where they exist, and prefer an OpenAI-compatible endpoint or open weights so the next migration is a configuration change.
Which One Should You Actually Use?
Voice agent that has to sound human: ElevenLabs v4 Turbo, first on the board at $0.04 per 1,000 characters.
Best quality per dollar: Speechify Simba 3.2, eighth on the board at $41 for our workload, if English is enough.
Already on OpenAI's speech endpoint: move before January 6, 2027; Inworld TTS-2 is the smallest code change.
Locked to a cloud contract: Azure Neural HD at $22, which ranks 40 places above Azure Neural.
Data that cannot leave the building: Kokoro under Apache 2.0, Piper if GPL-3.0 is acceptable, or Deepgram Flux TTS with a support contract.
Conclusion
The price spread is wider than the quality gap: Simba 3.2 ranks eighth in blind tests at about a tenth of Eleven v4's cost for our workload, Kokoro runs nine times faster than real time on a laptop CPU, and OpenAI is retiring its REST speech endpoint. Before you commit, hear your hardest line, normalized, on three candidates, from your own region.
Related DevToolLab Tools
- Text to Speech - hear a normalized script in your browser's voices before spending API credits.
- Character Counter - count the characters, spaces included, that per-character TTS pricing bills.
- Number to Words Converter - spell out amounts and ordinals the way you want them read.
- NATO Phonetic Alphabet Converter - turn confirmation codes into "Alpha Bravo" form a caller can write down.
Related Guides
- Best Speech-to-Text APIs - the recognition half of a voice pipeline, hosted and open models compared.
- Best AI Voice Agent Platforms - platforms that wire STT, an LLM and TTS into one call.
- Best GPU Cloud Providers - where to host open TTS models beyond a laptop.
