Back to all posts
Guide
12 min read

Best Text-to-Speech APIs in 2026, Compared

DevToolLab Team

DevToolLab Team

October 6, 2026

Best Text-to-Speech APIs in 2026, Compared

A text-to-speech API looks like the easy half of a voice product: send a string, get audio back. In production the bill is metered per character, every pause before the first word reads as lag, and the model decides on its own whether "Dr." means doctor or drive.

Three things moved in the past two weeks. On September 28, 2026, ElevenLabs shipped Eleven v4 at $0.08 per 1,000 characters. On October 1, OpenAI deprecated every model behind its speech endpoint, with a shutdown date of January 6, 2027. And on October 6, when we fed two popular open-source models the same appointment reminder, both read "Tue." as the number two.

How We Ranked and Priced Them

"Best" here means the Artificial Analysis Provider Voice Arena, a 92-model board where listeners pick the better of two clips without knowing which model made each, scored as Elo ratings as of October 6, 2026. A separate pronunciation test has reviewers check numbers, dates and codes.

Most APIs bill per character of input, so your own string is the meter; a few bill tokens, so their cost follows audio length. Every price below, checked on vendor pages on October 7, 2026, is for one workload: a US clinic's reminder line speaking 5 million characters a month, about 39,000 calls of this 128-character message:

"

Your appointment with Dr. Patel is confirmed for Tue., Oct. 13 at 2:30 PM. Your copay is $45.50. Questions? Call (555) 010-1234.

What Two Open Models Did With One Appointment Reminder

We ran the reminder through Kokoro 0.9.4 and Piper 1.8.0 on an M3 Pro MacBook Pro's CPU, then transcribed each WAV file back with faster-whisper 1.2.1 and its small.en model.

ModelWarm synthesisAudio lengthReal-time factorWhat came back
Kokoro-82M, af_heart voice1.52 s14.22 s0.107"confirmed for 2. Oct. 13", "Call 555-1012-34"
Piper, en_US-lessac-medium voice0.347 s13.66 s0.025"confirmed for 2.00. Oct. 13", "Call 555-010-1234"

Both models spoke "Tue." as the number two, and Kokoro regrouped the phone number into something a patient cannot dial. The fix was a dozen lines of Python that expand abbreviations and read phone numbers digit by digit before the text reaches any engine:

Python
import re

DAYS = {"Mon": "Monday", "Tue": "Tuesday", "Tues": "Tuesday", "Wed": "Wednesday", "Thu": "Thursday",
        "Thurs": "Thursday", "Fri": "Friday", "Sat": "Saturday", "Sun": "Sunday"}
MONTHS = {"Jan": "January", "Feb": "February", "Mar": "March", "Apr": "April", "Jun": "June",
          "Jul": "July", "Aug": "August", "Sep": "September", "Sept": "September", "Oct": "October",
          "Nov": "November", "Dec": "December"}

def alternation(words):
    return "|".join(sorted(words, key=len, reverse=True))

def normalize_for_tts(text: str) -> str:
    text = re.sub(rf"\b({alternation(DAYS)})\.,?", lambda m: DAYS[m[1]] + ",", text)
    text = re.sub(rf"\b({alternation(MONTHS)})\.", lambda m: MONTHS[m[1]], text)
    text = re.sub(r"\bDr\.(?= [A-Z])", "Doctor", text)
    # US phone numbers: one digit at a time, with a pause between the groups
    text = re.sub(r"\(?\b(\d{3})\)?[ .-]?(\d{3})[ .-](\d{4})\b",
                  lambda m: ", ".join(" ".join(group) for group in m.groups()), text)
    return text

if __name__ == "__main__":
    line = ("Your appointment with Dr. Patel is confirmed for Tue., Oct. 13 at 2:30 PM. "
            "Your copay is $45.50. Questions? Call (555) 010-1234.")
    print(normalize_for_tts(line))
text
Your appointment with Doctor Patel is confirmed for Tuesday, October 13 at 2:30 PM. Your copay is $45.50. Questions? Call 5 5 5, 0 1 0, 1 2 3 4.

With normalized input, both came back as "Tuesday, October 13" and "Call 555-010-1234", and hosted models need the same help: Eleven v4 scores 91.7 percent on the pronunciation test, Cartesia Sonic 3.6 74.5.

ElevenLabs Eleven v4 and v4 Turbo

ElevenLabs holds the top two spots: Eleven v4 Turbo at 1334 Elo and Eleven v4, launched September 28, 2026, at 1321.

ElevenLabs API pricing page listing Eleven v4 prices
ElevenLabs API pricing page listing Eleven v4 prices

Both cover 90-plus languages with voice cloning and lead the pronunciation test at 91.7 and 90.1 percent. v4 Turbo has a median inference latency of about 100 milliseconds, measured on ElevenLabs' servers rather than your network. The catch is price: Eleven v4 costs $400 for our workload and caps a request at 10,000 characters.

Pricing: Eleven v4 $0.08 per 1,000 characters · v4 Turbo and Flash v2.5 $0.04.

Alibaba Qwen-Audio TTS

Qwen-Audio-3.1-TTS-Plus from Alibaba ranks third at 1292 Elo, and Qwen-Audio-3.0-TTS-Plus ranks sixth at 1261.

Alibaba Cloud Model Studio documentation page for qwen-audio-3.0-tts-plus
Alibaba Cloud Model Studio documentation page for qwen-audio-3.0-tts-plus

Alibaba's docs say the Plus model improves low-resource languages and Chinese dialects. For US teams the limits are practical: Model Studio lists Beijing and Singapore deployments, with no US region, and the international pricing page publishes a rate for 3.0 but not yet for 3.1.

Pricing: Qwen-Audio-3.0-TTS-Plus $0.20 per 10,000 characters in Singapore · 3.0-TTS-Flash $0.15. Our workload: $100 on 3.0 Plus.

Cartesia Sonic-3.6

Cartesia ranks fourth at 1278 Elo with Sonic-3.6, a 44-language model it advertises at sub-90-millisecond latency.

Cartesia Sonic product page with a voice demo
Cartesia Sonic product page with a voice demo

Its best feature is versioning: sonic-3.6 tracks the latest stable snapshot, while sonic-3.6-YYYY-MM-DD pins one that never changes. It scored 74.5 percent on pronunciation despite claiming to read confirmation codes correctly, and plans cap concurrent requests at 2, 3, 5 and 15 from Free to Scale.

Pricing: Startup $49/month for 1.25 million · Scale $299 for 8 million · overage $38 to $65 per million. Our workload: $217.75 on Startup with overage.

Google, Amazon and Microsoft

The clouds are easy to buy on existing contracts, but their cheapest voices rank low.

Google Cloud Text-to-Speech product page
Google Cloud Text-to-Speech product page

Google Cloud Text-to-Speech has the fifth-ranked model, Gemini 3.8 Flash TTS at 1275 Elo and 89.5 percent on pronunciation, billed at $0.50 per million text tokens in and $9 per million audio tokens out through December 31, 2026, doubling on January 1, 2027. Chirp 3 HD voices, 58th, cost $30 per million characters with the first million free each month.

Amazon Polly charges $30 per million characters for Generative voices, 53rd, and $16 for Neural, 90th of 92. Azure AI Speech charges $22 for Neural HD, 28th, and $15 for Neural, 68th, in East US per Microsoft's Retail Prices API.

Inworld TTS-2

Inworld ranks seventh at 1251 Elo with Realtime TTS-2, and its cheaper TTS-2 Flash ranks twelfth at 1214.

Inworld Realtime TTS product page
Inworld Realtime TTS product page

Inworld serves OpenAI's POST /v1/audio/speech in OpenAI's own request format, so migrating means changing the SDK base URL to https://api.inworld.ai/v1 and the model to inworld-tts-2. Its docs list 200-plus languages.

Pricing: TTS-2 $25 per million characters on demand, $17.50 on Builder ($100/month) · TTS-2 Flash $15 on demand, $9 on Builder. Our workload: $125 on demand, $100 on Builder.

Speechify Simba 3.2

Speechify's developer API, sold separately from its reader app, ranks eighth with Simba 3.2 at 1242 Elo.

Speechify AI homepage for its text-to-speech and voice cloning API
Speechify AI homepage for its text-to-speech and voice cloning API

Its $10 per million characters on Starter, $8 on Pro and $6 on Scale is the lowest published per-character rate in the top ten, and Speechify cites independent first-audio times of 106 milliseconds on Coval and 123 on Voice Arena (September 24, 2026). The limit is language: self-serve accounts get English on Simba 3.2 and six languages on the older Simba 3.0.

Pricing: 500,000 characters free each month · Starter $10/month with 1.9 million included · Pro $99 with 13.5 million · Scale $499 with 78 million. Our workload: $41 on Starter.

MiniMax Speech 2.8

MiniMax ranks 19th with Speech 2.8 HD at 1173 Elo and 22nd with Speech 2.8 Turbo.

MiniMax Audio text-to-speech studio with voice cloning and voice design tools
MiniMax Audio text-to-speech studio with voice cloning and voice design tools

Its strengths are long-form audio and cloning: the asynchronous endpoint takes up to 1 million characters per request, and rapid voice cloning costs $1.50 per voice. It is also the most expensive API here per character.

Pricing: Speech 2.8 HD $100 per million characters · Speech 2.8 Turbo $60 · audio subscriptions from $5/month. Our workload: $500 on HD, $300 on Turbo.

Fish Audio S2.1 Pro

Fish Audio ranks 24th with S2.1 Pro at 1141 Elo.

Fish Audio homepage with a text-to-speech demo powered by S2.1 Pro
Fish Audio homepage with a text-to-speech demo powered by S2.1 Pro

The API has no subscription or monthly minimum, and concurrency rises from 5 to 15 to 50 requests as your total spend passes $100 and $1,000. It bills UTF-8 bytes rather than characters, so English text costs a per-character rate while accented or CJK text costs more.

Pricing: S2.1 Pro $15 per million UTF-8 bytes. Our workload: $75.

Deepgram, Hume, Rime and the Rest

Deepgram, absent from the board, launched Flux TTS on August 12, 2026: it keeps the whole call in context, reports what the caller heard when they interrupt, quotes time to first audio as low as 80 milliseconds and can run self-hosted. Aura-2 costs $0.030 per 1,000 characters and Flux TTS $0.045.

Deepgram text-to-speech product page
Deepgram text-to-speech product page

Hume Octave 2 ranks 61st, but Hume's pricing page now redirects to its homepage (checked October 7, 2026), so ask its sales team for a rate. Rime, whose Coda model ranks 55th, charges $0.05 per 1,000 characters for Coda and $0.03 for Mist v3, and sells on-premises deployments.

Luna TTS from VUI Labs (10th), Soniox (17th), Smallest.ai Lightning V3.1 Pro (20th) and Murf Falcon 2 (21st) also made the top 25, but we could not verify a per-character price on their pages; Soniox estimates about $0.70 per hour of generated speech.

OpenAI's Speech Endpoint Shuts Down January 6, 2027

OpenAI's deprecations page lists tts-1, tts-1-hd and both gpt-4o-mini-tts snapshots for shutdown on January 6, 2027, in a notice dated October 1, 2026, and names gpt-realtime-2.1-mini as the replacement.

OpenAI deprecations table for the January 6, 2027 TTS shutdown
OpenAI deprecations table for the January 6, 2027 TTS shutdown

That replacement runs on the Realtime API, a WebSocket or WebRTC session rather than a REST call, at $20 per million audio output tokens. Until the shutdown, tts-1, 41st, costs $15 per million characters. Do not start a new build on /v1/audio/speech; if you have one, Inworld's compatible endpoint is the smallest code change.

Open Models: Kokoro, Piper, Chatterbox and Breeze

Open weights trade the meter for setup work, and the license decides what you can ship.

Kokoro-82M model card on Hugging Face
Kokoro-82M model card on Hugging Face

Kokoro-82M ranks 54th and is Apache 2.0 licensed. Piper ran 40 times faster than real time on our CPU, but its MIT repository was archived on October 6, 2025, and development continues as piper1-gpl under GPL-3.0. Chatterbox from Resemble AI is MIT licensed, and its Multilingual V3 clones voices in 23-plus languages. Orpheus from Canopy Labs is Apache 2.0.

Archived rhasspy/piper repository on GitHub
Archived rhasspy/piper repository on GitHub

The best-ranked open weights, BreezeBlue's Breeze TTS 2 at 11th and Fish Audio's S2 Pro, are licensed for research and non-commercial use only. Kokoro and Piper also bundle espeak-ng, itself GPL-3.0; on macOS we had to point ESPEAK_DATA_PATH at a short path to its espeak-ng-data folder before either would run.

Side by Side

API or modelArena rankPer 1M characters5M characters/monthSelf-host
ElevenLabs v4 Turbo1$40$200No
ElevenLabs Eleven v42$80$400No
Cartesia Sonic-3.64$45 past Startup's 1.25M$217.75No
Google Gemini 3.8 Flash TTS5Token-billedToken-billedNo
Qwen-Audio-3.0-TTS-Plus6$20$100No
Inworld TTS-27$25$125No
Speechify Simba 3.28$6 to $10$41No
MiniMax Speech 2.8 HD19$100$500No
Fish Audio S2.1 Pro24$15 per M bytes$75No
Azure AI Speech Neural HD28$22$110No
OpenAI tts-141$15 until Jan 6, 2027$75No
Google Chirp 3 HD58$30, first 1M free$120No
Amazon Polly Neural90$16$80No
Deepgram Aura-2Not ranked$30$150No
Kokoro-82M54$0Hardware onlyApache 2.0

Ranks are Artificial Analysis Voice Arena positions out of 92 models on October 6, 2026; prices are list rates checked October 7, 2026.

How to Choose Without Migrating Twice

  1. Count your real characters. Export a month of the text you synthesize, SSML included, and multiply by the rates above.
  2. Normalize, then compare. Run your 50 hardest strings, such as order numbers and drug names, through normalize_for_tts and each candidate, then round-trip the audio through speech-to-text.
  3. Time the first audio byte from your region, us-east-1 for most US stacks, because vendor latency is measured on their side.
  4. Keep an exit. Pin dated model snapshots where they exist, and prefer an OpenAI-compatible endpoint or open weights so the next migration is a configuration change.

Which One Should You Actually Use?

Voice agent that has to sound human: ElevenLabs v4 Turbo, first on the board at $0.04 per 1,000 characters.

Best quality per dollar: Speechify Simba 3.2, eighth on the board at $41 for our workload, if English is enough.

Already on OpenAI's speech endpoint: move before January 6, 2027; Inworld TTS-2 is the smallest code change.

Locked to a cloud contract: Azure Neural HD at $22, which ranks 40 places above Azure Neural.

Data that cannot leave the building: Kokoro under Apache 2.0, Piper if GPL-3.0 is acceptable, or Deepgram Flux TTS with a support contract.

Conclusion

The price spread is wider than the quality gap: Simba 3.2 ranks eighth in blind tests at about a tenth of Eleven v4's cost for our workload, Kokoro runs nine times faster than real time on a laptop CPU, and OpenAI is retiring its REST speech endpoint. Before you commit, hear your hardest line, normalized, on three candidates, from your own region.

Related Posts

n8n vs Zapier vs Make: AI Automation 2026

n8n, Zapier and Make priced on one AI ticket-triage workflow: tasks vs credits vs executions, plus Activepieces, Windmill, Node-RED and OpenClaw.

By DevToolLab Team•

Best SQL Clients in 2026: 8 Tools Compared

DBeaver, DataGrip, Beekeeper Studio, TablePlus, DbGate, DbVisualizer, Chat2DB and VS Code's MSSQL extension, priced for a five-developer team.

By DevToolLab Team•

What Is an Agent Harness? Pi 1.0 Explained

An agent harness is the loop, tools, permissions and context around a model. Watch Pi 1.0 run one, then compare Claude Code, Codex, OpenCode and DeepSeek.

By DevToolLab Team•