tomai
Log in
☁️ Cloud · credits

Vocal Separator

Split any song into vocals and instrumental tracks with a neural stem model — instant karaoke

🪙 Neural stem separation on our server: 5 credits per job

🎙️ A neural model splits the audio into vocals + instrumental MP3 stems. Videos up to 10 minutes are supported.

Server-side FFmpeg handles any format — including HEVC/H.265 files your browser can't play. Files are deleted from the server within 30 minutes.

Stems for practice, remixes and karaoke

Split any track into vocals and instrumental — or four stems including drums and bass — delivered as a ZIP of separate files.

🎙️

AI vocal separation

A deep-learning model (Demucs) splits the track into voice and music by listening to the mix — not by filtering frequencies.

🎤

Vocals + instrumental

Get a karaoke instrumental and a clean acapella from one upload — remix, cover or study either stem.

📦

ZIP with both tracks

Both separated tracks come in one ZIP download, ready for your DAW or video editor.

Frequently asked questions

How good is the separation?

Demucs is currently the strongest open-source stem-splitting model. On clean studio mixes the vocals come out virtually instrumental-free. Extreme cases — heavy reverb, distorted vocals, dense metal mixes — show some bleed between stems, which is normal for any source-separation model.

Why do I get a ZIP instead of two separate downloads?

The two stems are generated together on our server, and a ZIP keeps them paired and preserves the file names. The ZIP contains exactly vocals.mp3 and instrumental.mp3.

What length limit is there?

10 minutes per job — neural stem separation is CPU-heavy, and longer inputs would exhaust our server capacity. Longer songs can be cut with the video trimmer first and split in parts.

Why does the instrumental still faintly contain vocals?

Separation is statistical, not perfect — frequencies overlapping between voice and instruments leave residue. The 4-stem model reduces it further than 2-stem; for fully clean instrumentals, pick tracks with sparse center-panned vocals.

Source separation models and outputs

Separation uses an SDR-optimized neural network executed on the worker CPU:

Each stem is encoded at the source's sample rate to MP3 320 kbps (or WAV if the input was lossless), bundled into a single ZIP. Songs up to 15 minutes are accepted. Bleed between stems is lowest on well-mastered pop/rock, higher on dense electronic mixes — the page says so upfront because honest expectations beat surprise refunds.

How it works

  1. 1

    Upload the song or video — up to 10 minutes long, any audio format or most videos.

  2. 2

    Start the job. The neural model analyzes the mix and separates voice from instrumentation.

  3. 3

    Download the ZIP containing vocals.mp3 and instrumental.mp3. CPU processing takes roughly as long as the song itself.

Related tools

What Is Vocal Separation?

A vocal separator splits a song into two clean tracks: the voice on its own and the music on its own. Rather than crude filtering that lets traces of the other half bleed through, a neural source-separation model listens to the whole mix and reconstructs the vocals and the instruments as separate audio streams. Upload an audio file — or a video with sound — up to 10 minutes long, and our server returns a ZIP containing vocals.mp3 and instrumental.mp3. Singers use it to rehearse without the melody, remixers isolate a hook or a beat, podcast editors strip unwanted speech from a clip, and karaoke fans get a backing track in seconds. The separation itself runs on our server with the open-source Demucs model.

What this tool can do

  • 🎤 Isolate the voice from any song as a clean acapella track
  • 🎹 Extract the instrumental for an instant karaoke version
  • 📦 Get both stems as MP3 files in a single ZIP download
  • 🧠 Separate stems with a neural model that understands the mix, not a simple filter
  • ⏱️ Accept audio and video files up to 10 minutes long
  • 🗑️ Delete uploads and results automatically within 30 minutes

When you would use it

  • 🎤 Practicing a song: sing along to the instrumental or study the isolated voice
  • 🎶 Hosting a karaoke night without hunting for minus-one versions
  • 🎧 Building remixes and mashups from clean stems
  • 🎙️ Removing the vocals from a podcast or video soundtrack
  • 🎓 Teaching music with the voice and accompaniment heard separately

Your file is sent to our VPS, where the Demucs neural network performs the separation, and both upload and results are deleted automatically within 30 minutes. A job costs credits, and running the model on the server's CPU takes about as long as the track itself plays. Each job is limited to 10 minutes of audio — longer tracks can be cut with a trimmer first. On clean studio mixes the stems come out virtually free of bleed, while heavily processed recordings, such as dense metal, heavy reverb or distorted vocals, may show some crosstalk — a normal property of any source-separation model.

Related tools