Turn any text into a narrated MP3 with subtitles
Turn a script into a narrated MP3 with neural voices, and get subtitles that line up with it to the millisecond.
The script travels to our own render endpoint, comes back as an MP3 with word-level timings, and is not stored anywhere. This is the engine that can be downloaded.
Enter starts a new paragraph, Shift+Enter keeps a line break inside one. Every paragraph can carry its own voice — that is how a dialogue gets narrated.
How it works
Bring in the script
Type it, paste it, or drop in a .txt, .md, .srt or .vtt. Each paragraph becomes an editable block. Nothing is synthesised on arrival.
Pick the voice
Search over 300 neural voices by language, locale or gender, and hear a sample line before committing the whole script to one.
Render it
Each paragraph is rendered on its own and the engine reports where every word falls. That is what feeds the waveform and the subtitles.
Take it away
MP3 and WAV of the exact audio you just heard, plus SRT and VTT subtitles built from the same timings.
Voices that sound like people
Over 300 neural voices across 140 locales, including multilingual ones that keep the same timbre when the script switches language.
Subtitles that really line up
The SRT and VTT files are built from the word and sentence timings the engine reports for that exact render — not estimated from a reading speed.
A voice per paragraph
Give each block its own voice and narrate a dialogue, an interview or an audiobook with several characters in a single pass.
Prosody you control
Speed, pitch and volume as a script-wide default, and overridable block by block whenever one line needs a different delivery.
Edit without starting over
Fix a typo in paragraph seven and only paragraph seven is rendered again. Everything already rendered stays exactly as it was.
An engine with no network
The device engine reads your script with the voices your operating system already has, with no request going anywhere. Listening only: browsers give no way to record it.
Narration and subtitles from one single render
Write or paste your script, choose from more than 300 neural voices, and get the finished audio together with SRT and VTT subtitles built from that very render. Every paragraph can take its own voice, its own speed and its own pitch, and editing one line only re-renders that line.
Where the audio is made, in plain words
With the device engine everything happens inside this tab, using the voices your system already has. With the neural engine the script is sent to our own endpoint, synthesised and sent straight back: no third-party relay in the middle, nothing written to a database, and no copy kept once the response is delivered. It is the one part that is not local, and it is the part that makes the download and the subtitles possible.
For narrating, studying and publishing
Voice-over a video and drop the matching subtitles onto the timeline. Turn a long article into audio for the commute. Give a course its narration with one voice for the lesson and another for the examples. Proofread your own writing by ear, which catches clumsy sentences no spell-checker ever will. And when the audio is done, send the subtitles on to the subtitle tool without downloading anything.
Your scripts, your history and your settings live in this browser and nowhere else. The neural engine receives only the paragraph it has to say out loud, and keeps nothing after answering. There is no account, no analytics on your text, and no log of what you had it read.
Frequently asked questions
QCan I download the audio, and is it the same as what I heard?+
Yes. The MP3 is the untouched stream the engine produced, stitched paragraph by paragraph with no re-encoding, so it is identical to what the player just played. The WAV is that same audio in an uncompressed form, for editors that prefer it.
QDo the subtitles really match the audio?+
They do, because they are not guessed. While it renders, the engine reports the exact instant every word and every sentence starts and how long it lasts. The SRT and VTT are built from those numbers, so they stay in sync even when the reading speed is changed.
QDoes my text leave the browser?+
With device voices, never. With neural voices, the paragraph being narrated is sent to our own render endpoint, which returns the audio and keeps nothing. That is the honest answer: neural voices cannot run inside a browser tab without downloading a model of hundreds of megabytes.
QIs there a length limit?+
Each paragraph is capped at 3,000 characters, which is a page of text; longer imports are split at sentence boundaries automatically. The script as a whole has no cap, though a very long one means many requests and takes proportionally longer.
QCan I use several voices in one script?+
Yes, and that is the point of splitting it into paragraphs. Open any block and give it its own voice, speed, pitch or volume. Everything you do not override keeps following the script-wide default.
QIf I fix a typo, does the whole thing render again?+
No. Each paragraph remembers the text and settings it was rendered with, so after an edit only the paragraphs that actually changed are sent again. On a long narration that is the difference between a few seconds and several minutes.
QWhy do device voices sound different on each computer?+
Because they belong to the operating system, not to us: Windows, macOS, Android, iOS and ChromeOS each ship their own set. Installing extra language packs adds voices to that list. Neural voices, by contrast, sound the same everywhere.