A 90-second explainer script runs about 225 spoken words, or roughly 1,300 characters. Characters, not minutes, are the unit you work in here. To make that voiceover: open the Text to Speech playground, paste the script, pick a voice, adjust the voice settings, and press Generate. That order matters, and so does the ranking behind it: voice choice shapes the result more than the model, and the model shapes it more than any slider. Modern systems reach this quality because neural architectures now synthesise speech straight from text, with Tacotron 2 scoring 4.53 on mean opinion score against 4.58 for professionally recorded speech. The rest of this guide covers what those four steps leave out.
Script to Finished Audio in One Sitting
Most first-time users get a usable track in under twenty minutes. The time goes into voice auditioning, not generation.
The four moves that matter
- Paste or type your script into the input box.
- Select a voice from your voice list.
- Adjust settings only if the default read is wrong.
- Generate, listen, and re-roll the weak lines.
What to have ready first
A finished script, a target length, and a reference clip of a voice you like. Write the script the way a person speaks. Read it aloud once. Cut any sentence you stumble on, because the model will stumble in the same place. Mark names and acronyms you expect to be mangled, along with any number that must be read a specific way. Decide the target length in characters, since that is what your plan meters. A rough conversion helps: 150 spoken words per minute lands near 870 characters per minute of finished audio.
Setting Up: Account, Character Budget, and the Right Editor
Signup is email or Google, and you land in the creative platform immediately.
Test before you commit budget
Free access is usually enough to audition voices and produce short clips. Paid tiers add character volume and clearer commercial terms. Plan structures and rates change often, so read the vendor’s own pricing page on the day you build the budget rather than trusting any figure quoted in an article, including this one.
Text to Speech versus the longer-form editors
The playground is for single takes: one script, one voice, one file. The platform also ships longer-form editors for multi-speaker projects and for audio timed against video. Start in the playground. Move up only when you need several speakers or frame-accurate timing. The jump costs you an afternoon of learning the timeline, so avoid it for a one-voice explainer.
API access if you are batching
If you plan to generate hundreds of clips, the API is the route rather than the interface. Programmatic generation is usually metered differently from manual use, so check the developer pricing page before you commit to a batch workflow. For a single script, the interface is faster.
Voices, Models, and the Settings That Change the Read
This is where quality is won or lost. Voice selection has the biggest impact on your output, ahead of model choice, ahead of every slider you can drag.
Four ways to get a voice
| Voice source | How you get it | Best for | Main caveat |
|---|---|---|---|
| Default voices | Curated, ready in the picker | First drafts, quick tests | Widely used, so it can sound familiar |
| Voice Library | Browse and add to your list | Matching a specific accent or age | Quality varies by source recording |
| Voice Design | Describe the voice in a prompt | Characters, unusual timbres | Synthetic, needs more auditioning |
| Voice cloning | Instant or professional cloning | A consistent brand voice | Requires consent and clean source audio |
Voices are not equal, because much depends on the recordings used to build them. A voice trained on clean studio audio holds together across a long read; one trained on thin source material wobbles. Audition three before committing, and audition them on your hardest sentence rather than your easiest one.
Accents travel, and that surprises people
Any voice can technically read any language, but it keeps its original accent. Take an English voice and ask it for French: you get correct French with an English accent. For a native read, pick a native voice from the library instead, or clone a speaker who already speaks that language. This single choice fixes more localisation complaints than any slider does.
The sliders, in plain terms
Stability controls how much the delivery varies between generations. Lower stability gives more emotion and more inconsistency. Similarity controls how tightly the output tracks the original voice. Style exaggeration pushes expressiveness. Speaker boost sharpens resemblance. Speed adjusts pace. Change one at a time. Moving three sliders together makes it impossible to tell which one improved the read, and you will burn characters proving nothing.
Tuning by content type
- Long narration and audiobooks: higher stability, low style. Consistency across an hour beats a great single line.
- Ads and social clips: lower stability, more style. You want energy in a nine-second read.
- Course modules and tutorials: mid stability, slightly slower speed, so a listener can take notes.
- Character work: prompt-designed voices, high style, and expect several re-rolls.
Building Your First Track: A Worked Walkthrough
Here is a concrete run, with numbers you can copy.
The brief
A hardware startup needs a 45-second store-shelf demo. The script is 112 words, 663 characters with spaces. Target: one female narrator, warm, mid-Atlantic accent, no music.
What actually happened
Voice auditioning took eleven minutes across five library voices, generating the same two opening sentences each time. Generation of the full script took under thirty seconds. Two lines came back wrong: the product name read as three syllables instead of four, and one question landed flat.
The fix and the final cost
The name was respelled phonetically in the input text and regenerated. The flat question was fixed by dropping stability and re-rolling twice. Total characters consumed: 663 for the main take, plus roughly 900 across auditions and re-rolls. Call it 1,563 characters for a finished 45-second asset, or about 35 characters per second of audio. Use that ratio to estimate any project: minutes of audio times 60, times 35, plus a third for auditions. A 12-minute course module therefore lands near 33,600 characters once re-rolls are counted, which is the number to check against your plan before you promise a deadline. Compared with booking a studio session for the same 45 seconds, the saving here was mostly calendar time: no scheduling, no pickup session two days later when marketing changed one word.
When It Goes Wrong: Fixes and Traps
Almost every problem falls into four buckets.
Names and acronyms come out wrong
Respell the word phonetically inside the script. “Zimran” becomes “Zim ran”. Say acronyms as separated letters. Regenerate only that line rather than the whole file, then splice it back in your editor.
The delivery drifts across a long file
Raise stability and split the script into shorter blocks. Long unbroken input gives the model more room to wander in tone. Blocks of 300 to 500 characters are easy to review and cheap to regenerate.
The read sounds flat or robotic
That is usually punctuation, not the model. Add commas where you would breathe. Replace semicolons with full stops. Break any sentence over 28 words into two. Punctuation is the only prosody control you get inside the text itself, so use it deliberately.
Traps that cost people a re-record
- Writing for the eye instead of the ear, then blaming the voice.
- Auditioning voices on generic filler text rather than your real opening lines.
- Locking a voice before checking it holds up on numbers and dates.
- Cloning a voice you do not have written permission to use.
- Exporting a compressed file, then discovering the video editor needs lossless audio.
- Forgetting that re-rolls consume your character budget too.
Product, Course, App and Platform Experience: Where the Audio Ships
Generating audio is the easy half. Landing it in a real product is where teams get stuck.
Video, ads, and social
Short-form video is the most common destination. Export, drop into the timeline, and expect to nudge timing by a few frames. Punchy reads beat perfect reads at this length. Listeners scrolling on a phone forgive a rushed syllable; they do not forgive a dull opening line.
Courses and internal training
Narrated course modules benefit most, because a script update means regenerating one paragraph instead of rebooking a studio. Keep a master script file so versions stay traceable, with the voice name and settings recorded next to each module. Six months later, nobody remembers which slider produced the good take. Whoever signs off on a module should understand the tool well enough to judge the output, not just generate it: a reviewer who knows how the model handles emphasis, pacing and pronunciation will catch the takes that sound almost right. If you are building the underlying AI skills alongside the tooling, Explore Coursiv AI lessons for structured practice.
Podcasts, audiobooks, and accessibility
Synthetic speech started life as assistive technology for readers with visual impairments and dyslexia, and that remains a legitimate primary use rather than an afterthought. Small teams also use it to keep marketing audio in-house; the Small Business Administration’s marketing guidance is a sober starting point for deciding which channels deserve narrated content at all.
Decision Framework: What to Know Before Deciding on an AI Voice
Not every project should use a synthetic voice. Score yours before you commit.
Five questions, honestly answered
- Does the script change often? Frequent edits favour synthetic voices heavily.
- How long is the finished audio? Under two minutes, a human read is still cheap.
- Does the audience need to trust a specific person? If yes, use that person.
- Do you need the same voice for years? Then plan for a cloned or licensed voice.
- Is emotional range the point? Grief, comedy, and sarcasm remain hard, and a listener notices the miss instantly.
Three or more answers favouring synthesis means go ahead. Two or fewer means hire a narrator and keep the tool for drafts.
Rights, consent, and disclosure
Clone only voices you have documented permission to use, in writing, with the scope spelled out. In the United States, copyright protects original works of authorship such as literary and dramatic works, and your script counts, while copyright itself is the body of law covering those works. It does not protect facts, ideas, or methods, and rights in a person’s voice are a separate question governed by contract and state law. Some platforms and clients require disclosure of synthetic audio. Ask before delivery, not after. Written approval of the voice, the script, and the disclosure wording takes ten minutes and prevents an expensive argument.
Honest caveats
Output quality is uneven across languages and across voices built from weak source recordings. Vendor documentation proves that a feature exists, not that it performs well for your accent or subject matter. Numbers, dates, and technical jargon still need spot-checking by ear. And every claim about plan limits should be re-checked on the vendor’s own pricing page, because those change without notice.
You will probably want how to make money with ai voiceovers soon, and how to use suno ai to make songs shortly after.