Blog · 5 min read

Why noise removal can't remove a second voice

A cleanup tool hears every voice as speech worth keeping. Extraction keeps one person. The difference, and a free way to use it.

By · Published Aug 24, 2026 · Updated Sep 9, 2026

You have a recording, and it has too many people in it. An interview where the host keeps finishing the guest’s sentences. A customer call with a colleague talking in the background. A voice memo that caught the whole kitchen. The words you need are in there; they are tangled up with everyone else’s.

Can you remove one person’s voice from a recording?

Yes. Give a tool a few seconds of the voice you want to keep, and it can return that voice alone, with every other speaker and the background removed. The technique is called target speaker extraction, and there is a free version of it in the browser, no account needed. The rest of this guide shows how to use it and what to expect.

If you have searched this before, you have probably read that it cannot be done. Forum answers compare it to removing an egg from a cake mixture, and for the tools they had in mind, they were right. Equalizers, phase tricks, and vocal removers all work on frequencies, and two voices share the same ones. The same goes for the effects in your editor, which we go through one by one here. Extraction works on identity instead: it learns what one speaker sounds like and keeps only that.

The other fix people reach for, a noise remover, fails for a different reason. That reason is worth two paragraphs, because it tells you what to use instead.

Why noise removal keeps the second voice

A cleanup tool answers one question: which parts of this signal are speech, and which parts are not? The hum, the traffic, the keyboard, the music bed: all of that goes. Every voice stays. That is not a flaw. Protecting speech is the entire design.

So the second person in the room survives every pass. To a noise remover they are not noise; they are exactly the thing it was built to keep. Run a two-voice recording through cleanup and you get the same two voices back, cleaner than before.

What works: target speaker extraction

The tool you actually want answers a different question: which parts of this signal are this person? Give it a short sample of the voice you care about. It builds a voiceprint from that sample, follows the voice through the recording, and returns it alone. Every other voice, and everything else in the scene, is treated the way a hum would be: as interference, removed.

The technical name is target speaker extraction. It is one of three related techniques that get mixed up constantly; diarization labels who spoke when, and speech separation splits everyone onto unlabeled tracks. We wrote a field guide to all three if you want the map. What matters here is the shape of the contract: one sample in, one speaker out.

Where to run it

We ship a free tool that does this in the browser: Isolate a Speaker. Upload the recording, mark three clean seconds of the voice you want, and the result comes back beside the original. Recordings up to 60 seconds, no account. The walkthrough with screenshots, before-and-after clips, and the steps for putting the result back into Premiere, Resolve, CapCut, iMovie or Final Cut is in how to remove one person’s voice from a video.

If the voice you get back is the right person but the room is still around them, that is a second, separate job. Extraction decides which voice; speech enhancement decides how it sounds. Run them in that order, and see why your recording sounds like a bathroom for the second half.

What a good result sounds like

A few things about the output are worth knowing in advance.

  • The gaps are correct. When your speaker is not talking, the right output is silence, not a quieter version of everyone else. Empty stretches mean the tool can tell the difference between your speaker pausing and your speaker being absent.
  • The voice can sound a little narrower. The output is tuned for speech, and the before and after are matched, so the difference you hear between them is the separation rather than a change in recording quality.
  • The sample decides the run. Almost every disappointing result traces back to the sample: too short, or not actually alone. The fix is a cleaner three seconds, not a different recording.

When the recording is really a stream

The batch tool works on files, after the fact. The harder version of the problem is live: a voice agent that answers whoever the microphone picks up, a contact-center pipeline transcribing the television, hardware that hears the room instead of its owner.

That is the problem Crele is building for: a real-time speaker extraction engine, the same contract at conversation speed, coming soon. If the file you keep cleaning up is really a stream, join the waitlist.

Hear it first

The engine is coming soon. Join the waitlist for early access, and we'll write only when there's something real to try.

What would you use it for?