tomai
Log in
☁️ Cloud · credits

Text to Speech

Turn written text into natural MP3 speech with neural voices — Mandarin plus two English voices

🪙 Neural speech synthesis on our server: 2 credits per job

Up to 20,000 characters per job — roughly 30 minutes of speech. The MP3 downloads when the job finishes.

Type or paste the text you want spoken — up to 20,000 characters per job 0 / 20000

Voice engine and practical limits

Synthesis runs on a neural TTS engine with four voice profiles:

Long texts are split on sentence boundaries into chunks synthesized independently, then concatenated gaplessly — so a 15-minute narration doesn't degrade toward the end. Output is MP3 (48 kHz) delivered like any cloud job. Punctuation controls pacing: em-dashes create breaths, ellipses slow the cadence. SSML is intentionally not exposed — plain text keeps results predictable.

Natural voices sized for real projects

Turn articles into narrations up to 20,000 characters per job, with distinct male and female voices for Chinese and English.

🗣️

Natural neural voices

A neural text-to-speech model reads your text with natural rhythm and intonation — clearer than robotic synth voices.

🌍

Mandarin + English voices

Choose from a Mandarin voice and two English voices, so both Chinese and international content can be voiced.

🎵

Instant MP3 download

The spoken audio is converted to MP3 and ready to download right away — for voiceovers, study audio or accessibility.

How it works

  1. 1

    Type or paste your text — up to 20,000 characters per job, roughly 30 minutes of speech.

  2. 2

    Pick a voice: Mandarin Chinese, US English female, or US English male.

  3. 3

    Start the job and download your MP3. The server synthesizes it with a neural voice model — usually just a few seconds.

Frequently asked questions

Is this real neural TTS or old robotic speech synthesis?

Real neural TTS. The voices are Piper models trained on thousands of hours of studio recordings — pauses, intonation and stress sound natural, nothing like the robot voices of early synthesis.

Which voices are available?

Three: Mandarin Chinese, US English female (Amy) and US English male (Joe). The voice list lives server-side and can grow later.

Is there a text length limit?

Yes — 20,000 characters per job (about 30 minutes of audio). Longer text can be split into several jobs and stitched together with the audio merge tool.

Can I use the audio commercially?

Yes — narrations you generate from text you have rights to are yours to publish, including monetized videos and courses. Attributing the voice isn't required by our terms.

Related tools

What Is Text to Speech?

Text to Speech turns written words into spoken audio. Type or paste any text and the tool returns a natural-sounding MP3 recording, synthesized by a neural voice model (Piper) running on our server. It is the opposite of audio-to-text transcription: instead of reading a recording, you write the words and get the voice back. Three voices are built in — Mandarin Chinese, US English female and US English male — and because there is no file to upload, the text itself is the request. The job usually finishes in a few seconds. Video creators use it for narration, students for pronunciation practice, and busy readers for listening to articles on the go.

What this tool can do

  • 📝 Convert typed or pasted text into a spoken MP3 recording
  • 🗣️ Choose from three neural voices: Mandarin Chinese, US English female or US English male
  • 📏 Handle up to 20,000 characters per job — roughly 20 to 30 minutes of speech
  • ⚡ Synthesize in near real time, usually within a few seconds
  • 🎧 Play the result right on the page and download the MP3 file
  • 🔤 Track your input with a live character counter — no file upload needed at all

When you would use it

  • Creating narration or voiceover for a video, slideshow or presentation
  • Practicing pronunciation of English or Chinese by listening to a clear voice
  • Turning an article or document into audio to listen to on the move
  • Producing voice for e-learning lessons, announcements or podcast episodes
  • Reading long text aloud when your eyes need a break

Your text travels over an encrypted connection to our VPS server, where the neural engine synthesizes the audio, and the job files are deleted after processing — nothing is kept. The service consumes credits per job. The current limits: three built-in voices (one Mandarin, two English), a 20,000-character cap per job, and MP3 output at a fixed natural pace — custom voices, extra languages and speed control are not available yet.

Related tools