oLoveTools
Text to speech

Turn any text into a narrated MP3 with subtitles

Turn a script into a narrated MP3 with neural voices, and get subtitles that line up with it to the millisecond.

300+ neural voices in 140 localesMP3 and WAV of the exact take you heardSRT and VTT from the engine's own timings

The script travels to our own render endpoint, comes back as an MP3 with word-level timings, and is not stored anywhere. This is the engine that can be downloaded.

Voice
Reading speed1.00×
Voice pitch1.00×
Volume100%
Script0 words · 0 characters
0

Enter starts a new paragraph, Shift+Enter keeps a line break inside one. Every paragraph can carry its own voice — that is how a dialogue gets narrated.

How it works

1

Bring in the script

Type it, paste it, or drop in a .txt, .md, .srt or .vtt. Each paragraph becomes an editable block. Nothing is synthesised on arrival.

2

Pick the voice

Search over 300 neural voices by language, locale or gender, and hear a sample line before committing the whole script to one.

3

Render it

Each paragraph is rendered on its own and the engine reports where every word falls. That is what feeds the waveform and the subtitles.

4

Take it away

MP3 and WAV of the exact audio you just heard, plus SRT and VTT subtitles built from the same timings.

Voices that sound like people

Over 300 neural voices across 140 locales, including multilingual ones that keep the same timbre when the script switches language.

Subtitles that really line up

The SRT and VTT files are built from the word and sentence timings the engine reports for that exact render — not estimated from a reading speed.

A voice per paragraph

Give each block its own voice and narrate a dialogue, an interview or an audiobook with several characters in a single pass.

Prosody you control

Speed, pitch and volume as a script-wide default, and overridable block by block whenever one line needs a different delivery.

Edit without starting over

Fix a typo in paragraph seven and only paragraph seven is rendered again. Everything already rendered stays exactly as it was.

An engine with no network

The device engine reads your script with the voices your operating system already has, with no request going anywhere. Listening only: browsers give no way to record it.

text to speech

Narration and subtitles from one single render

Write or paste your script, choose from more than 300 neural voices, and get the finished audio together with SRT and VTT subtitles built from that very render. Every paragraph can take its own voice, its own speed and its own pitch, and editing one line only re-renders that line.

300+ neural voices in 140 locales
MP3 and WAV of the exact take you heard
SRT and VTT from the engine's own timings
A different voice for each paragraph
Offline device voices for instant listening
No account, no watermark, no limits per day

Where the audio is made, in plain words

With the device engine everything happens inside this tab, using the voices your system already has. With the neural engine the script is sent to our own endpoint, synthesised and sent straight back: no third-party relay in the middle, nothing written to a database, and no copy kept once the response is delivered. It is the one part that is not local, and it is the part that makes the download and the subtitles possible.

For narrating, studying and publishing

Voice-over a video and drop the matching subtitles onto the timeline. Turn a long article into audio for the commute. Give a course its narration with one voice for the lesson and another for the examples. Proofread your own writing by ear, which catches clumsy sentences no spell-checker ever will. And when the audio is done, send the subtitles on to the subtitle tool without downloading anything.

What we do and do not keep

Your scripts, your history and your settings live in this browser and nowhere else. The neural engine receives only the paragraph it has to say out loud, and keeps nothing after answering. There is no account, no analytics on your text, and no log of what you had it read.

Frequently asked questions

QCan I download the audio, and is it the same as what I heard?+

Yes. The MP3 is the untouched stream the engine produced, stitched paragraph by paragraph with no re-encoding, so it is identical to what the player just played. The WAV is that same audio in an uncompressed form, for editors that prefer it.

QDo the subtitles really match the audio?+

They do, because they are not guessed. While it renders, the engine reports the exact instant every word and every sentence starts and how long it lasts. The SRT and VTT are built from those numbers, so they stay in sync even when the reading speed is changed.

QDoes my text leave the browser?+

With device voices, never. With neural voices, the paragraph being narrated is sent to our own render endpoint, which returns the audio and keeps nothing. That is the honest answer: neural voices cannot run inside a browser tab without downloading a model of hundreds of megabytes.

QIs there a length limit?+

Each paragraph is capped at 3,000 characters, which is a page of text; longer imports are split at sentence boundaries automatically. The script as a whole has no cap, though a very long one means many requests and takes proportionally longer.

QCan I use several voices in one script?+

Yes, and that is the point of splitting it into paragraphs. Open any block and give it its own voice, speed, pitch or volume. Everything you do not override keeps following the script-wide default.

QIf I fix a typo, does the whole thing render again?+

No. Each paragraph remembers the text and settings it was rendered with, so after an edit only the paragraphs that actually changed are sent again. On a long narration that is the difference between a few seconds and several minutes.

QWhy do device voices sound different on each computer?+

Because they belong to the operating system, not to us: Windows, macOS, Android, iOS and ChromeOS each ship their own set. Installing extra language packs adds voices to that list. Neural voices, by contrast, sound the same everywhere.

Related searches

text to speechfree TTStext to MP3neural voicesAI voice overgenerate subtitlestext to SRTnarrate an articleaudiobook voicespeech synthesis
Part of the oLoveTools suite
oLoveTools

Free, private and precise text-to-speech, with subtitles included.

Where the audio is made, in plain words

With the device engine everything happens inside this tab, using the voices your system already has. With the neural engine the script is sent to our own endpoint, synthesised and sent straight back: no third-party relay in the middle, nothing written to a database, and no copy kept once the response is delivered. It is the one part that is not local, and it is the part that makes the download and the subtitles possible.

For narrating, studying and publishing

Voice-over a video and drop the matching subtitles onto the timeline. Turn a long article into audio for the commute. Give a course its narration with one voice for the lesson and another for the examples. Proofread your own writing by ear, which catches clumsy sentences no spell-checker ever will. And when the audio is done, send the subtitles on to the subtitle tool without downloading anything.

Narration and subtitles from one single render

Your scripts, your history and your settings live in this browser and nowhere else. The neural engine receives only the paragraph it has to say out loud, and keeps nothing after answering. There is no account, no analytics on your text, and no log of what you had it read.

text to speechfree TTStext to MP3neural voicesAI voice overgenerate subtitlestext to SRTnarrate an articleaudiobook voicespeech synthesis

Frequently asked questions

Can I download the audio, and is it the same as what I heard?

Yes. The MP3 is the untouched stream the engine produced, stitched paragraph by paragraph with no re-encoding, so it is identical to what the player just played. The WAV is that same audio in an uncompressed form, for editors that prefer it.

Do the subtitles really match the audio?

They do, because they are not guessed. While it renders, the engine reports the exact instant every word and every sentence starts and how long it lasts. The SRT and VTT are built from those numbers, so they stay in sync even when the reading speed is changed.

Does my text leave the browser?

With device voices, never. With neural voices, the paragraph being narrated is sent to our own render endpoint, which returns the audio and keeps nothing. That is the honest answer: neural voices cannot run inside a browser tab without downloading a model of hundreds of megabytes.

Is there a length limit?

Each paragraph is capped at 3,000 characters, which is a page of text; longer imports are split at sentence boundaries automatically. The script as a whole has no cap, though a very long one means many requests and takes proportionally longer.

Can I use several voices in one script?

Yes, and that is the point of splitting it into paragraphs. Open any block and give it its own voice, speed, pitch or volume. Everything you do not override keeps following the script-wide default.

If I fix a typo, does the whole thing render again?

No. Each paragraph remembers the text and settings it was rendered with, so after an edit only the paragraphs that actually changed are sent again. On a long narration that is the difference between a few seconds and several minutes.

Why do device voices sound different on each computer?

Because they belong to the operating system, not to us: Windows, macOS, Android, iOS and ChromeOS each ship their own set. Installing extra language packs adds voices to that list. Neural voices, by contrast, sound the same everywhere.

© 2026 oLoveToolsAbout

Next step

Convert the audio AudioSnap

TTSBolt | Free Text-to-Speech with MP3, WAV and Subtitle Export

Narrate any text with 300+ neural voices, download the MP3 or WAV, and export SRT and VTT subtitles generated from the render itself. A different voice per paragraph, free and with no account.

Frequently asked questions

Can I download the audio, and is it the same as what I heard?

Yes. The MP3 is the untouched stream the engine produced, stitched paragraph by paragraph with no re-encoding, so it is identical to what the player just played. The WAV is that same audio in an uncompressed form, for editors that prefer it.

Do the subtitles really match the audio?

They do, because they are not guessed. While it renders, the engine reports the exact instant every word and every sentence starts and how long it lasts. The SRT and VTT are built from those numbers, so they stay in sync even when the reading speed is changed.

Does my text leave the browser?

With device voices, never. With neural voices, the paragraph being narrated is sent to our own render endpoint, which returns the audio and keeps nothing. That is the honest answer: neural voices cannot run inside a browser tab without downloading a model of hundreds of megabytes.

Is there a length limit?

Each paragraph is capped at 3,000 characters, which is a page of text; longer imports are split at sentence boundaries automatically. The script as a whole has no cap, though a very long one means many requests and takes proportionally longer.

Can I use several voices in one script?

Yes, and that is the point of splitting it into paragraphs. Open any block and give it its own voice, speed, pitch or volume. Everything you do not override keeps following the script-wide default.

If I fix a typo, does the whole thing render again?

No. Each paragraph remembers the text and settings it was rendered with, so after an edit only the paragraphs that actually changed are sent again. On a long narration that is the difference between a few seconds and several minutes.

Why do device voices sound different on each computer?

Because they belong to the operating system, not to us: Windows, macOS, Android, iOS and ChromeOS each ship their own set. Installing extra language packs adds voices to that list. Neural voices, by contrast, sound the same everywhere.

Related searches

text to speech, free TTS, text to MP3, neural voices, AI voice over, generate subtitles, text to SRT, narrate an article, audiobook voice, speech synthesis