Scripts and the reads they produced
Speech has no thumbnail. Each entry is the text that was pasted and the waveform of what came back.
- Prompt
“Every generation you run lands in your library the moment it finishes, tagged and searchable, with the settings that produced it kept alongside the file. Nothing is filed by hand and nothing has to be renamed later. Come back a month from now, search for the brief rather than the filename, and the take is where you left it — along with the version before it, and the one before that. When a script changes, you do not start again. Open the run that is closest, change the line that moved, and render it against the same voice and the same settings as the take it replaces. That is the part that saves the afternoon: not the first generation, but the fourteenth, when the difference between them is one word and everything else has to match.”
0:45 - Prompt
“I didn't think it would work. I want to be honest about that, because I argued against trying it twice, and I was fairly sure I was right. We had a version of this three years ago and it was embarrassing. It sounded like a machine reading a hostage note. So when they asked me to listen to it again, I said yes the way you say yes to a friend showing you a magic trick — politely, already composing the thing I would say afterwards. And then it worked. Not almost. Not close enough to fix in the edit. It read the line the way I would have read it, including the pause I would have taken, which nobody wrote down anywhere.”
0:39 - Prompt
“The Starter plan gives you 60,000 credits a month, which works out at roughly 1,000 images or about 12 videos. Pro raises that to 150,000 — call it 2,700 images, or 30 videos. Ultimate starts at 260,000 and scales up through 475,000 to 700,000, which is around 12,700 images or 150 videos at the top of the range. Credits refresh monthly and unused credits do not roll over, so there is no advantage in hoarding them. You can top up at any point in the month. Text to speech is charged per 1,000-character block with a one-block floor, which means a 40-character line and a 900-character script cost exactly the same: one block, at the rate of whichever model you picked.”
0:55
A script, a voice, and the delivery
Paste your script
Up to 5000 characters in one run, enforced as you type. These models read punctuation as direction: a comma is a breath and a full stop is a beat — write the pauses in rather than asking for them.
Pick a model, then a voice
Ten models across four families. ElevenLabs brings twenty named voices, Minimax seventeen, Qwen nine and Cartesia fifty, which you can filter by nationality and accent.
Tune the read and generate
Speed on every family, plus stability, similarity and style on ElevenLabs, pitch and emotion on Minimax, or volume, emotion and language on Cartesia. The audio lands in your workspace library, and failed runs are refunded.
What Text to Speech takes, and what it returns
What Text to Speech needs from you
- Script
- Text, up to 5000 characters per run, enforced as you type
- Voice
- One of 96 named voices — 20 on ElevenLabs, 17 on Minimax, 9 on Qwen and 50 on Cartesia
- Price
- Charged per 1000-character block, with a one-block floor — a short line costs a full block
- Languages
- Qwen3 TTS covers eleven with automatic detection; ElevenLabs Multilingual v2 holds one voice across thirty
Export specifications
- Per run
- Up to 5000 characters of script
- Models
- Ten, across ElevenLabs, Minimax, Qwen and Cartesia
- Voices
- 96 named voices across four families
- Price
- Per 1000-character block, at the rate of the model you chose, with a one-block floor
- Failed runs
- Refunded
- Delivery
- Rendered audio file in your workspace library
Not just which voice — how it reads the line
The controls differ by model, because they are the model's own parameters rather than a common wrapper. Speed is the one every family offers.
-
Script
5000 characters per run
Enforced as you type, and priced per 1000-character block with a one-block floor. Longer scripts split into runs — break at a paragraph so the phrasing does not fall apart at the seam.
-
ElevenLabs voices
20 named voices, Aria the default
Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily and Bill.
-
Minimax voices
17 named voices, Wise Woman the default
Wise Woman, Friendly Person, Inspirational Girl, Deep Voice Man, Calm Woman, Casual Guy, Lively Girl, Patient Man, Young Knight, Determined Man, Lovely Girl, Decent Boy, Imposing Manner, Elegant Man, Abbess, Sweet Girl 2 and Exuberant Girl.
-
Qwen voices
9 named voices, Vivian the default
Vivian, Serena, Uncle Fu, Dylan, Eric, Ryan, Aiden, Ono Anna and Sohee. Cartesia has its own fifty, across nine countries, with filters for nationality and accent.
-
Speed
0.7–1.2 on ElevenLabs and Qwen; 0.6–1.5 on Cartesia; 0.5–2 on Minimax
The one control every family offers. The narrow band is narrow on purpose — pushed further, the read stops sounding like a person. Minimax gives the wider range, for anything from a slow read to a compressed disclaimer.
-
Stability
0–1, default 0.5
ElevenLabs only. Low is expressive and varies between takes; high is consistent and flatter.
-
Similarity
0–1, default 0.75
ElevenLabs only. How closely the render holds to the reference voice.
-
Style
0–1, default 0
ElevenLabs only. Pushes the delivery away from neutral. Zero is the default for a reason.
-
Pitch
−12 to +12, default 0
Minimax only. Semitone shift on the voice, without changing the pace.
-
Emotion
Neutral, happy, sad, angry, fearful, disgusted, surprised
Minimax only. Seven named settings, chosen rather than coaxed out of the punctuation.
-
Language
Eleven on Qwen3 TTS, with automatic detection
Auto, English, Chinese, Spanish, French, German, Italian, Japanese, Korean, Portuguese and Russian.
Engine choice
Pick the one that matches your source.
-
ElevenLabs v3
ExpressiveThe most expressive. Reads emotion out of the writing.
Narration that has to carry feeling
-
ElevenLabs Multilingual v2
MultilingualThirty languages, one voice across all of them.
Scripts that switch language
-
ElevenLabs Turbo v2.5
CheaperHalf the price and most of the quality. Good for drafts.
Drafts and volume
-
Minimax Speech 2.8 Turbo
ControlsEmotion, pitch and volume as real controls rather than hints.
Directed performance, quickly
-
Minimax Speech 2.8 HD
FidelityThe same voices, rendered at higher fidelity.
A finished directed read
-
Qwen3 TTS 1.7B
LanguagesEleven languages with automatic detection.
Non-English scripts
-
Qwen3 TTS 0.6B
FastThe smaller, faster one. Same voices.
Iterating on a non-English script
-
Cartesia Sonic 3.6
InstantCartesia's newest. Finishes while you wait, with fifty voices in forty-four languages.
A read you need now
-
Cartesia Sonic 3.5
InstantFinishes while you wait — nothing to come back for. The same fifty voices.
A read you need now
-
Minimax Voice Clone
CloningReads your script in a voice from a recording you supply.
A voice of your own
Built for anyone who needs a script read aloud
-
Video creators
Narrate a cut without booking a booth, and re-render the line when the script changes.
-
Course and training teams
Keep one named voice across dozens of modules, including the ones written months apart.
-
Product teams
Voice a demo or an in-app walkthrough, in more than one language, from the same script.
Built for Growth at Every Stage
Every plan includes every model and every feature. Plans only change how many credits you get and how many generations run at once.
- Credits refresh monthly
- Top-up additional credits anytime
- Unused credits don't roll over
Not ready for the commitment?
FAQ
Frequently asked questions
Everything you need to know about Text to Speech on BeHooked.
Which voices can I choose from?
Ninety-six named voices across four families. ElevenLabs has twenty: Aria, which is the default, plus Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily and Bill. Minimax has seventeen: Wise Woman, the default, plus Friendly Person, Inspirational Girl, Deep Voice Man, Calm Woman, Casual Guy, Lively Girl, Patient Man, Young Knight, Determined Man, Lovely Girl, Decent Boy, Imposing Manner, Elegant Man, Abbess, Sweet Girl 2 and Exuberant Girl. Qwen has nine: Vivian, the default, plus Serena, Uncle Fu, Dylan, Eric, Ryan, Aiden, Ono Anna and Sohee. Cartesia has fifty across nine countries — American, British, Indian, Australian, Irish, South African, Canadian, New Zealand and Singaporean, plus four that speak Hindi — with filters for nationality and accent.
How many models are there, and which should I pick?
Ten. ElevenLabs v3 is the most expressive and reads emotion out of the writing; Multilingual v2 holds one voice across thirty languages; Turbo v2.5 is half the price and most of the quality, which makes it the draft option. Minimax Speech 2.8 Turbo gives you emotion, pitch and volume as real controls rather than hints, and 2.8 HD is the same voices at higher fidelity. Qwen3 TTS 1.7B covers eleven languages with automatic detection and 0.6B is the smaller, faster one with the same voices. Cartesia Sonic 3.6 and 3.5 finish while you wait, in forty-four languages. Minimax Voice Clone reads your script in a voice from a recording you supply.
How is it priced?
Per 1000-character block, at the rate of the model you choose, with a one-block floor — a single short line still costs one block, so it is worth batching short lines into one run where the script allows. ElevenLabs Turbo v2.5 is the cheapest of the ElevenLabs three at about half the price. Failed runs are refunded.
Can I change the pace?
Yes, on every family. Speed runs from 0.7 to 1.2 on ElevenLabs and Qwen, from 0.6 to 1.5 on Cartesia, and from 0.5 to 2 on Minimax. The narrow band is narrow deliberately: pushed further, the read stops sounding like a person. Before reaching for it, note that these models read punctuation as direction — a comma is a breath and a full stop is a beat, so write the pauses in rather than asking for them.
What does stability actually do?
It is an ElevenLabs control running from 0 to 1, defaulting to 0.5. Low values give a more expressive read that varies between takes; high values give a consistent, flatter one. Similarity, at 0.75 by default, is the separate control for how closely the render holds to the reference voice, and style, at 0 by default, pushes the delivery away from neutral. None of the three appears on the other families.
Can I make it sound happy, or angry?
On Minimax, yes — it takes one of seven named emotions: neutral, happy, sad, angry, fearful, disgusted or surprised, alongside a pitch shift of −12 to +12 semitones. Both are Minimax-only. On ElevenLabs the equivalent handle is the style control, from 0 to 1.
How long a script can I run at once?
Up to 5000 characters, enforced as you type. Longer scripts need splitting into runs — break at a paragraph rather than mid-sentence so the phrasing holds across the seam.
What languages does it cover?
Qwen3 TTS covers eleven with automatic detection: Auto, English, Chinese, Spanish, French, German, Italian, Japanese, Korean, Portuguese and Russian. ElevenLabs Multilingual v2 is the other multilingual option, holding one voice across thirty languages.