Make a voice recording
sound studio-clean
Upload an audio or video file of a person talking. The voice comes back clean with noise and the room removed. Free up to 60 seconds with no account.
Drop an audio or video file here
WAV, MP3, M4A, FLAC, MP4 or MOV. We use the first 60 seconds.
If it sounds like you recorded in a bathroom
Hollow, boomy, tinny, far away, like a tin can: those are all one problem, and it is the room rather than the microphone. Hard walls hand every word back a few milliseconds late, and the recording keeps the room along with the voice.
This is how you clean up a voice recording: the background noise goes, the room goes, and the same words come back in the speaker's own voice.
It regenerates rather than filters, which is how it reaches what a filter cannot. Reverb is your own voice arriving late, in the same frequencies as the voice you want, so there is no band to subtract. It is also why a result can sound a touch different from the take. The technical name is speech enhancement, and taking the room off is dereverberation.
It is built for one voice. A second person in the recording is not something it is made for, and what comes back for them is not a promise we make. If you want one of them on their own, isolate that speaker first and enhance the track that comes back.
A noise remover keeps every voice and drops the noise. Isolation keeps one voice and drops the others. This one keeps the voice and hands the recording back rebuilt.
Questions
- Why does my recording sound like I recorded in a bathroom?
- Because the room is in the recording. Hard walls, tile and glass send every word back at the microphone a few milliseconds late, and those reflections pile up into the hollow, ringing sound people call a bathroom. It is reverb, and it is on the recording rather than on your voice.
- Why does my recording sound hollow, tinny or far away?
- They are one problem heard from three angles: too much room and not enough direct voice. Hollow and tin can are reverb. Far away is a microphone too distant to win against it. Tinny is a thin recording with the body of the voice missing. Rebuilding the voice addresses all three.
- What is speech enhancement?
- Taking a recording of speech and returning it cleaner and easier to understand than it arrived: noise gone, room gone, the voice brought forward. Classical versions filter the signal. Current models rebuild the voice out of what they can still recognise in it, which is what lets reverb come off at all.
- Can AI remove reverb better than a filter?
- Yes, and reverb is the clearest case for it. A filter works on frequency, and reverb is your own voice arriving late, sitting in the same frequencies as the voice you want to keep. There is no band to cut. A model that rebuilds the voice simply leaves the reflections out.
- How do I remove echo in Audacity, Premiere or CapCut?
- Mostly you cannot. Noise reduction and gates in an editor work on level and frequency, which is why they duck steady hiss but leave a room ringing, and why a hard gate makes speech sound clipped instead. Taking a room off needs a model that knows what a voice is. That is what runs here, in the browser, with no install.
- How do I fix a recording made in a bad room?
- Re-record it closer to the microphone if you still can, because no tool beats a better take. If the take is all you have, run it through this one: the room comes off, and what comes back is the voice without it. Free up to 60 seconds, no account.
- Does it remove echo and room sound?
- Yes. Reverb is the part a filter struggles with most, and rebuilding the voice leaves the room out entirely.
- How is this different from a noise remover?
- A noise remover subtracts noise from the recording. This rebuilds the voice, so what you keep is the voice rather than the recording minus its noise.
- What does it do to the voice?
- It removes the background noise and the room from a recording of one person and returns the voice clean. Same words, same speaker.
- Does AI enhancement change your voice?
- It lands close to the take without being identical. The model rebuilds your voice rather than filtering it, so the timbre can shift a little, most audibly on a rough take where there is less left to work from. The words and the speaker stay yours.
- Why does AI-enhanced audio sound robotic?
- Two different faults get the same word. A subtractive tool cuts noise out of the recording and leaves holes and warbling where it cut. A regenerative one rebuilds the voice, so its failure is the opposite: too smooth, or a word landing slightly wrong. Neither is text to speech, which has no recording underneath it at all. Which fault you have, and what fixes it
- Does it remove other people's voices?
- No. It is built for one voice at a time. To take one speaker out of a crowd, use the speaker isolation tool first, then process the track it returns. How to remove one person's voice from a video
- Why does the result sound a little different?
- The voice is rebuilt rather than filtered, so it lands close to the take without being identical. The before and after are matched in the player, so what you hear when you switch is the enhancement.
- Why does the result sound narrower than my upload?
- The output is tuned for speech at 16 kHz, so it can sound narrower than a full-band original. The before and after are matched to the same bandwidth.
- Does it work on phone calls, voice memos and meeting recordings?
- Yes, wherever one person is talking. A call recording or a voice memo is exactly the kind of recording it handles. A meeting with several people in it is not: isolate the speaker you want first, then enhance the track that comes back. WAV, MP3, M4A, FLAC, MP4 and MOV all work.
- Can I upload a video?
- Yes. Drop it on the tool, MP4 or MOV. Your browser reads the audio track out of the file, so only the audio is sent, and the voice comes back enhanced as an audio file to lay back over the picture.
- Can it take music, or a podcast with two hosts?
- It is built for one voice. Music behind the voice is treated as noise and removed. For two hosts, isolate the one you want first.
- What if it changed a word, or it does not sound like me?
- Give the result a thumbs down and say why. Then try a take with less noise: the harder the recording, the more the model has to fill in.
- Is it really free?
- Yes, free for recordings up to 60 seconds, no account needed. If you need to process longer files, join our waitlist.
- What happens to my audio?
- The result link stops working after an hour. For more on how we handle your files, see the privacy page.
- Is this the real-time engine?
- No, this is the free batch tool for recorded files. The real-time engine is coming soon. If you want early access to it, join our waitlist.