Turn any text or PDF into natural speech.
AI text-to-speech for scripts, articles, and books. Paste text or upload a PDF, pick a voice, get an mp3 in seconds.
Free: 10,000 characters a month. No card required.
Your audio
Take your audio where your listeners already are





Hear the voices
76+ voices across 17 languages, many with expressive styles. Press play, or type your own text in the live demo.
Ready to try them? Start free →
How it works
From text to audio in three steps.
Paste text or upload a PDF
Up to 100,000 characters. Drop in a script, an article, or a whole chapter.
Pick a neural voice
Dozens of lifelike MAI Voice 2 voices across accents and languages, including HD.
Get your audio
An mp3 in seconds for short text, a minute or two for long-form — play it back instantly.
The models behind the voices
Three state-of-the-art neural TTS engines. Pick the right one for the job.
MAI Voice 2
DefaultMicrosoft's flagship neural TTS. Preset voices across 15+ languages and 18 locales, many with 18 expressive styles each — cheerful, whispering, sad, angry, and more.
Benchmark
Fooled 45% of listeners in a blind human-vs-AI test (Alphasignal, 2025) — closest to human of any TTS tested.
Gemini 3.1 Flash
Google's multilingual TTS. 30 preset voices with auto-detected language coverage across 70+ languages. Inline mood tags like [whispers] and [laughs] drive prompt-driven emotion.
Best for
Long-tail languages MAI doesn't cover (Hindi, Vietnamese, Swahili) and per-tag emotion control.
Qwen Audio TTS Plus
Alibaba's multilingual TTS via DashScope. Multilingual auto-detect like Gemini, with low-latency raw PCM output wrapped to WAV.
Best for
Asian-language coverage and fast PCM streaming where WAV is preferred over MP3.
Who it’s for
Wherever you need to turn words into sound.
Podcasters & creators
Draft voiceovers, intros, and narration without a mic.
Students
Turn notes and readings into audio to revise on the go.
Accessibility
Listen to any text — helpful for dyslexia and low vision.
Language learners
Hear natural pronunciation across 15+ languages.
How tts.audio compares
Side-by-side with the major AI text-to-speech services. Pricing verified July 2026.
| Feature | tts.audio | ElevenLabs | Murf | Play.ht | Speechify |
|---|---|---|---|---|---|
| Cheapest paid | $0.13/day · 100K chars | $6/mo · 30K credits | $19/mo · ~30K chars | $39/mo · 600K words | $29/mo (reader app) |
| Free tier | 10,000 chars / month | 10,000 credits / month | 10 min gen · no download | 5,000 words / month | 10 robotic voices |
| TTS engine | MAI Voice 2 + Gemini + Qwen | Proprietary v2 / v3 | Proprietary | Proprietary | Proprietary |
| PDF → audio upload | Yes — native | No | No | No | Yes (reader) |
| Languages | 15+ (MAI) + 70+ multilingual | 32 | 20+ | 29 | 60+ |
| Voice cloning | No | Yes (Instant + Professional) | Yes (paid) | Yes (paid) | Yes (Studio add-on) |
| Commercial use | Pro ($4) and up | Starter ($6) and up | Creator ($19) and up | Paid plans | Studio (separate) |
| Audio output | MP3 · WAV | MP3 · 44.1kHz PCM on Pro+ | MP3 · WAV | MP3 · WAV | MP3 |
Sources: each vendor’s public pricing page, accessed 2026-07-25.tts.audio is character-based SaaS.
Simple, character-based pricing
Pay for what you generate. Upgrade or cancel anytime.
Creator
$0.63/day
500,000 characters / month
- All voices + PDF upload
- Commercial use
| Compare plans | Free | Pro | Creator | Studio |
|---|---|---|---|---|
| Price | $0 | $0.13/day | $0.63/day | $2.63/day |
| Characters / month | 10,000 | 100,000 | 500,000 | 2,000,000 |
| Text input | Yes | Yes | Yes | Yes |
| PDF upload | Yes | Yes | Yes | Yes |
| Voice styles | Yes | Yes | Yes | Yes |
| Priority queue | — | — | — | Yes |
| Commercial use | — | Yes | Yes | Yes |
No long-term contracts. Cancel anytime. The free plan never needs a card.
Frequently asked questions
Everything you need to know about AI text-to-speech with tts.audio.
Reviewed by Antonio Foti, Founder. Updated July 2026.
What is AI text-to-speech?
AI text-to-speech (TTS) is technology that converts written text into spoken audio using neural networks. Modern AI voices like Microsoft MAI Voice 2 produce natural-sounding speech with realistic intonation, pacing, and emotion across 15 languages and 18 locales (Microsoft, 2025), and fooled 45% of listeners in a blind human-versus-AI test reported by Alphasignal in 2025 — far beyond the robotic output of older systems.
How does tts.audio convert a PDF into audio?
Upload a PDF and tts.audio extracts its text, splits it into chunks at sentence boundaries, synthesizes each chunk with a neural voice, then concatenates the audio into a single MP3 you can play or download. Long documents up to 100,000 characters are handled by a background worker in about a minute or two.
How much text can I convert for free?
The free plan includes 10,000 characters per month with no credit card required. Paid plans raise the limit: Pro gives 100,000 characters for $0.13/day, Creator gives 500,000 characters for $0.63/day, and Studio gives 2,000,000 characters for $2.63/day with priority processing. Every paid plan includes commercial usage rights.
What voices and languages are available?
tts.audio uses Microsoft MAI Voice 2 neural voices, offering 76+ lifelike voices across 15+ languages and 18 locales (Microsoft, 2025), many with multiple expressive styles (cheerful, sad, excited, whispering, and more). Gemini 3.1 Flash and Qwen Audio voices extend coverage to 70+ languages via auto-detection. You pick a voice and optional style before generating, and the same setup is used consistently across short and long-form conversions.
Can I use the generated audio commercially?
Commercial usage rights are included with every paid plan (Pro, Creator, and Studio). The free tier is licensed for personal use only. Generated MP3 files are available to play and download for 24 hours via a secure link.
How long are generated audio files kept?
Each generated MP3 is available through a secure link for 24 hours, after which it is automatically removed from storage. You can generate it again any time within your monthly character limit if you need a fresh copy.
Ready to turn your text into speech?
Paste text or upload a PDF and get your first audio in seconds.
No card required.
