A 90-second explainer script runs about 225 spoken words, or roughly 1,300 characters. Characters, not minutes, are the unit you work in here. To make that voiceover: open the Text to Speech playground, paste the script, pick a voice, adjust the voice settings, and press Generate. That order matters, and so does the ranking behind it: voice choice shapes the result more than the model, and the model shapes it more than any slider. Modern systems reach this quality because neural architectures now synthesise speech straight from text, with Tacotron 2 scoring 4.53 on mean opinion score against 4.58 for professionally recorded speech. The rest of this guide covers what those four steps leave out.

Script to Finished Audio in One Sitting

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Most first-time users get a usable track in under twenty minutes. The time goes into voice auditioning, not generation.

The four moves that matter

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
  1. Paste or type your script into the input box.
  2. Select a voice from your voice list.
  3. Adjust settings only if the default read is wrong.
  4. Generate, listen, and re-roll the weak lines.

What to have ready first

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

A finished script, a target length, and a reference clip of a voice you like. Write the script the way a person speaks. Read it aloud once. Cut any sentence you stumble on, because the model will stumble in the same place. Mark names and acronyms you expect to be mangled, along with any number that must be read a specific way. Decide the target length in characters, since that is what your plan meters. A rough conversion helps: 150 spoken words per minute lands near 870 characters per minute of finished audio.

Setting Up: Account, Character Budget, and the Right Editor

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Signup is email or Google, and you land in the creative platform immediately.

Test before you commit budget

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Free access is usually enough to audition voices and produce short clips. Paid tiers add character volume and clearer commercial terms. Plan structures and rates change often, so read the vendor’s own pricing page on the day you build the budget rather than trusting any figure quoted in an article, including this one.

Text to Speech versus the longer-form editors

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

The playground is for single takes: one script, one voice, one file. The platform also ships longer-form editors for multi-speaker projects and for audio timed against video. Start in the playground. Move up only when you need several speakers or frame-accurate timing. The jump costs you an afternoon of learning the timeline, so avoid it for a one-voice explainer.

API access if you are batching

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

If you plan to generate hundreds of clips, the API is the route rather than the interface. Programmatic generation is usually metered differently from manual use, so check the developer pricing page before you commit to a batch workflow. For a single script, the interface is faster.

Voices, Models, and the Settings That Change the Read

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

This is where quality is won or lost. Voice selection has the biggest impact on your output, ahead of model choice, ahead of every slider you can drag.

Four ways to get a voice

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
Voice sourceHow you get itBest forMain caveat
Default voicesCurated, ready in the pickerFirst drafts, quick testsWidely used, so it can sound familiar
Voice LibraryBrowse and add to your listMatching a specific accent or ageQuality varies by source recording
Voice DesignDescribe the voice in a promptCharacters, unusual timbresSynthetic, needs more auditioning
Voice cloningInstant or professional cloningA consistent brand voiceRequires consent and clean source audio

Voices are not equal, because much depends on the recordings used to build them. A voice trained on clean studio audio holds together across a long read; one trained on thin source material wobbles. Audition three before committing, and audition them on your hardest sentence rather than your easiest one.

Accents travel, and that surprises people

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Any voice can technically read any language, but it keeps its original accent. Take an English voice and ask it for French: you get correct French with an English accent. For a native read, pick a native voice from the library instead, or clone a speaker who already speaks that language. This single choice fixes more localisation complaints than any slider does.

The sliders, in plain terms

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Stability controls how much the delivery varies between generations. Lower stability gives more emotion and more inconsistency. Similarity controls how tightly the output tracks the original voice. Style exaggeration pushes expressiveness. Speaker boost sharpens resemblance. Speed adjusts pace. Change one at a time. Moving three sliders together makes it impossible to tell which one improved the read, and you will burn characters proving nothing.

Tuning by content type

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
  • Long narration and audiobooks: higher stability, low style. Consistency across an hour beats a great single line.
  • Ads and social clips: lower stability, more style. You want energy in a nine-second read.
  • Course modules and tutorials: mid stability, slightly slower speed, so a listener can take notes.
  • Character work: prompt-designed voices, high style, and expect several re-rolls.

Building Your First Track: A Worked Walkthrough

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Here is a concrete run, with numbers you can copy.

The brief

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

A hardware startup needs a 45-second store-shelf demo. The script is 112 words, 663 characters with spaces. Target: one female narrator, warm, mid-Atlantic accent, no music.

What actually happened

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Voice auditioning took eleven minutes across five library voices, generating the same two opening sentences each time. Generation of the full script took under thirty seconds. Two lines came back wrong: the product name read as three syllables instead of four, and one question landed flat.

The fix and the final cost

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

The name was respelled phonetically in the input text and regenerated. The flat question was fixed by dropping stability and re-rolling twice. Total characters consumed: 663 for the main take, plus roughly 900 across auditions and re-rolls. Call it 1,563 characters for a finished 45-second asset, or about 35 characters per second of audio. Use that ratio to estimate any project: minutes of audio times 60, times 35, plus a third for auditions. A 12-minute course module therefore lands near 33,600 characters once re-rolls are counted, which is the number to check against your plan before you promise a deadline. Compared with booking a studio session for the same 45 seconds, the saving here was mostly calendar time: no scheduling, no pickup session two days later when marketing changed one word.

When It Goes Wrong: Fixes and Traps

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Almost every problem falls into four buckets.

Names and acronyms come out wrong

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Respell the word phonetically inside the script. “Zimran” becomes “Zim ran”. Say acronyms as separated letters. Regenerate only that line rather than the whole file, then splice it back in your editor.

The delivery drifts across a long file

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Raise stability and split the script into shorter blocks. Long unbroken input gives the model more room to wander in tone. Blocks of 300 to 500 characters are easy to review and cheap to regenerate.

The read sounds flat or robotic

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

That is usually punctuation, not the model. Add commas where you would breathe. Replace semicolons with full stops. Break any sentence over 28 words into two. Punctuation is the only prosody control you get inside the text itself, so use it deliberately.

Traps that cost people a re-record

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
  • Writing for the eye instead of the ear, then blaming the voice.
  • Auditioning voices on generic filler text rather than your real opening lines.
  • Locking a voice before checking it holds up on numbers and dates.
  • Cloning a voice you do not have written permission to use.
  • Exporting a compressed file, then discovering the video editor needs lossless audio.
  • Forgetting that re-rolls consume your character budget too.

Product, Course, App and Platform Experience: Where the Audio Ships

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Generating audio is the easy half. Landing it in a real product is where teams get stuck.

Video, ads, and social

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Short-form video is the most common destination. Export, drop into the timeline, and expect to nudge timing by a few frames. Punchy reads beat perfect reads at this length. Listeners scrolling on a phone forgive a rushed syllable; they do not forgive a dull opening line.

Courses and internal training

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Narrated course modules benefit most, because a script update means regenerating one paragraph instead of rebooking a studio. Keep a master script file so versions stay traceable, with the voice name and settings recorded next to each module. Six months later, nobody remembers which slider produced the good take. Whoever signs off on a module should understand the tool well enough to judge the output, not just generate it: a reviewer who knows how the model handles emphasis, pacing and pronunciation will catch the takes that sound almost right. If you are building the underlying AI skills alongside the tooling, Explore Coursiv AI lessons for structured practice.

Podcasts, audiobooks, and accessibility

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Synthetic speech started life as assistive technology for readers with visual impairments and dyslexia, and that remains a legitimate primary use rather than an afterthought. Small teams also use it to keep marketing audio in-house; the Small Business Administration’s marketing guidance is a sober starting point for deciding which channels deserve narrated content at all.

Decision Framework: What to Know Before Deciding on an AI Voice

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Not every project should use a synthetic voice. Score yours before you commit.

Five questions, honestly answered

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
  1. Does the script change often? Frequent edits favour synthetic voices heavily.
  2. How long is the finished audio? Under two minutes, a human read is still cheap.
  3. Does the audience need to trust a specific person? If yes, use that person.
  4. Do you need the same voice for years? Then plan for a cloned or licensed voice.
  5. Is emotional range the point? Grief, comedy, and sarcasm remain hard, and a listener notices the miss instantly.

Three or more answers favouring synthesis means go ahead. Two or fewer means hire a narrator and keep the tool for drafts.

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Clone only voices you have documented permission to use, in writing, with the scope spelled out. In the United States, copyright protects original works of authorship such as literary and dramatic works, and your script counts, while copyright itself is the body of law covering those works. It does not protect facts, ideas, or methods, and rights in a person’s voice are a separate question governed by contract and state law. Some platforms and clients require disclosure of synthetic audio. Ask before delivery, not after. Written approval of the voice, the script, and the disclosure wording takes ten minutes and prevents an expensive argument.

Honest caveats

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Output quality is uneven across languages and across voices built from weak source recordings. Vendor documentation proves that a feature exists, not that it performs well for your accent or subject matter. Numbers, dates, and technical jargon still need spot-checking by ear. And every claim about plan limits should be re-checked on the vendor’s own pricing page, because those change without notice.

You will probably want how to make money with ai voiceovers soon, and how to use suno ai to make songs shortly after.

Frequently asked questions

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
What is ElevenLabs, in one sentence?
It is an AI audio platform whose creative tooling spans text to speech, voice cloning, dubbing, music, and sound effects, aimed at creators who need finished audio rather than raw research models.
How do I create my own voice?
Two paths. Voice Design generates a voice from a written prompt. Voice cloning builds one from recorded audio, in instant or professional form, with the professional route needing more source material and consent. Either way, keep the original recordings; if the clone disappoints you, better source audio is the usual fix.
What languages can I generate?
Coverage spans dozens of languages, and the current list sits on the vendor’s own product page. Accent quality still depends on picking a native voice rather than forcing an English voice into another language.
Why does the same script sound different each time?
Low stability settings introduce deliberate variation between generations. Raise stability for repeatable reads, and regenerate individual lines instead of whole files when only one moment is off.