Blog · 7 min read

How to separate two voices in Audacity or Premiere

Every audio editor has a button that looks like it should do this. Here is what each one actually does to two voices in one track.

By · Published Aug 25, 2026 · Updated Sep 9, 2026

You recorded two people on one microphone. Now you need one of them on their own. So you open your editor, find the effect whose name sounds exactly right, run it, and get back either the same two voices as before or a hollow version of both.

Can you separate two voices recorded on one microphone?

Yes, but almost certainly not with the effect you were about to try. The technique that keeps one chosen voice is called target speaker extraction, and there is a free version of it in the browser. The vocal isolation tools in audio editors were built for music, and the noise tools were built to protect speech, so both of them keep the person you wanted gone. One editor now ships a genuine speaker splitter as a paid feature.

Here is the whole landscape in one table, then the reasoning behind each row.

EffectBuilt forTwo voices, one mic
Audacity, Vocal Reduction and Isolationcenter-panned vocals in stereo musicNo. Needs true stereo, and removes everything centered
Audacity, Noise Reductionsteady background noiseNo. Designed to protect speech, so both voices survive
Audacity, Music Separation (OpenVINO)drums, bass, vocals, otherNo. Both people are “vocals” and land on one stem
Audacity, Noise Suppression (OpenVINO)speech against noiseNo. Same reason. Speech is the thing it keeps
Premiere, Enhance Speechclarity, noise, room toneNo. Built around a single primary speaker
Premiere, Separate Overlapping Speakerssplitting crosstalk onto tracksYes. Paid, billed per second of audio
Audition, Noise Reduction with a voice printreducing a captured profilePartly. Lowers a voice, does not remove it
Target speaker extractionkeeping one enrolled voiceYes. Sample the voice you want, get it alone

Why Audacity cannot separate two voices

Audacity’s Vocal Reduction and Isolation is the effect everyone tries first, and it is not a voice tool at all. It is a stereo trick. Music mixes usually place the lead vocal equally in the left and right channels, so inverting one channel against the other cancels whatever sits in the middle.

Two constraints follow, and Audacity’s own manual states both. The first is that the input must be a true stereo track and not mere (dual) mono. A phone recording, a voice memo, or anything captured on one microphone is mono, so there is no left-against-right difference to work with and the effect has nothing to do. The second is the more fundamental one: the plug-in “does not know what kind of audio is in the center. All is removed or isolated equally.”

That sentence is the entire problem. Two people talking into one microphone are both in the center. The effect cannot prefer one of them, because it has no concept of a person. It only knows position, and they share a position.

Noise Reduction fails for the opposite reason. You give it a sample of noise, and it subtracts that profile from the rest. It is built for hum, traffic, air conditioning: sounds that hold still. A human voice does not hold still, and every voice in the file is the thing the effect was written to preserve.

Do Audacity’s AI plugins separate speakers?

No. Since version 3.5 Audacity has offered a free set of Intel OpenVINO AI plugins, and two of them look like the answer. Neither is.

Music Separation runs a stem-splitting model that returns four tracks: drums, bass, vocals, and other instruments. It is genuinely good at that job. But it splits by instrument family, not by person, so two people talking are both vocals. They come out mixed together on the same stem, which is where you started.

Noise Suppression runs speech-enhancement models. Their entire purpose is to tell speech apart from everything that is not speech, and to keep the speech. Feed it two voices and it protects both of them, cleanly.

The pattern holds across every tool on this page. An effect can only separate things it can tell apart, and none of these can tell one voice from another. They sort by stereo position, by instrument, or by speech against noise. Two people in one room are identical on all three axes.

Can Premiere Pro separate overlapping speakers?

Yes. Premiere now has a feature called Separate Overlapping Speakers, also labelled Separate Crosstalk, which analyses a clip and generates a separate track per speaker while leaving the original intact. It is the one editor feature on this list that does what people have been asking editors to do for years.

Three things to know before you rely on it. It is a premium feature, priced in generative credits at one credit per second of audio at the time of writing. It accepts clips from one second to ten minutes, so a long recording has to be trimmed to the section that matters. And it adds a two second handle to each side, which counts toward the cost.

Premiere’s other audio feature, Enhance Speech, is not this. It cleans up dialogue by reducing noise and room reflection, and Adobe’s documentation is explicit that it is built around one primary speaker and does poorly on overlapping voices. It makes both of your speakers sound better. It does not remove either.

That is a different job, and a real one: if your problem is a single voice in a bad room rather than two voices in one track, see why your recording sounds like a bathroom.

What about Adobe Audition?

Audition has no speaker separation feature, and the standing answer in its own community forums is that this cannot be done: with both voices on one microphone there is nothing to tell them apart. That answer was correct for every tool Audition ships.

The nearest thing is Noise Reduction with a captured voice print, where you sample the unwanted voice and apply it as a profile. It pulls that voice down in level. It does not take it out, and it takes some of your speaker with it, because the two voices share most of their frequency content.

What actually works: extraction, not subtraction

Everything above tries to subtract something. The reason they fail is that subtraction needs a property that separates the two voices, and at the level these effects operate on, there isn’t one.

Extraction inverts the question. Instead of describing what to remove, you point at the voice you want to keep. Give the model a few seconds of that person talking alone. It builds a voiceprint from the sample, follows that voice through the recording, and returns it by itself. Every other speaker is treated the way a hum would be.

This is a different contract from what Premiere’s crosstalk tool offers, and the difference decides which one you want:

  • Separation takes a mixture and returns every speaker on their own track. You get all of them, then pick yours. It works best when the number of speakers is small and known.
  • Extraction takes a mixture plus a sample and returns one speaker. You name who you want up front. The number of other people in the room does not change the job, because they are all just interference.

If you want one person out of a crowded scene, or the recording has more voices in it than you want to sort through, extraction is the shorter path. Our field guide to diarization, separation and extraction covers the full picture, including why transcription tools mislabel speakers.

How to keep one voice, with no editor

We ship a free browser tool that does the extraction: Isolate a Speaker. Recordings up to 60 seconds, no account, nothing to install.

  1. Upload the audio or video file. WAV, MP3, M4A and FLAC all work, and so do MP4 and MOV: drop the clip straight from your timeline export and your browser reads the audio track out of it, so only the audio is sent. Mono is fine. Mono is the normal case here.
  2. Mark three seconds of the voice you want. Find a moment where that person is talking alone and select it. Clean matters more than long. If the other voice is inside your sample, the result keeps both.
  3. Run it, and compare. The isolated track comes back beside the original, so the difference you hear between them is the extraction itself.

Sixty seconds is the cap on the free tool, so this is not the route for a full podcast episode. It is the route for the passage that matters, and it is the fastest way to hear whether extraction solves your problem at all. If the recording is a video, how to remove one person’s voice from a video walks through it with screenshots, and covers putting the result back in your editor.

When the recording is really a stream

Every tool on this page works on a file, after the fact. The harder version of the problem is live: a voice agent that answers whoever the microphone picks up, a contact-center pipeline that transcribes the television, hardware that hears the room instead of its owner. No editor effect helps there, because there is no file to open yet.

That is what Crele is building: a real-time speaker extraction engine, the same one sample in, one speaker out contract at conversation speed, coming soon. If the file you keep cleaning up is really a stream, join the waitlist.

Hear it first

The engine is coming soon. Join the waitlist for early access, and we'll write only when there's something real to try.

What would you use it for?