Your microphone was six inches from your mouth. The recording sounds like you were across the room, in a stairwell, talking into a tube. Nothing was wrong with the microphone, and turning the gain down will not fix it.
Why does my recording sound like a bathroom?
Because your microphone heard the room as well as you. Every hard surface sent a copy of your voice back, hundreds of times a second, each one slightly late. Those copies pile up behind the words as a wash. The effect is called reverb, and a small tiled or empty room produces a great deal of it.
This is the single most common reason a recording sounds wrong when the gear was fine. It is also why the problem gets worse the further you sit from the microphone: the direct sound falls away with distance, the reflected sound does not, so the balance tips toward the room.
Hollow, tinny, boomy, far away: mostly the same problem
People describe this sound in a dozen ways, and the descriptions are more useful than they look. Most of them point at reverb. A few point somewhere else entirely, and it is worth knowing which is which before you go looking for a fix.
| What you hear | What it usually is | Fixable afterwards |
|---|---|---|
| Like a bathroom, a tube, a stairwell | Reverb. Hard surfaces, empty room | Yes |
| Hollow, distant, far away | Reverb, plus too much distance from the mic | Yes |
| Boomy or roomy | Reverb weighted to low frequencies, small room | Yes |
| Tinny, like a tin can | Often reverb, sometimes a phone or laptop mic’s own tone | Usually |
| Echo you can count, a distinct repeat | True echo. A large space, or a duplicated track | Sometimes |
| Muffled, like a blanket over it | Not reverb. Obstruction, or the mic facing away | Partly |
| Crackle, dropouts, robotic stutter | Not reverb. A connection or codec problem | Rarely |
Two rows deserve a note. Doubled audio is not reverb at all: it is the same take playing twice, slightly offset, usually because two sources were recorded onto one timeline. Look for a second track before you reach for a tool. And a muffled recording is the opposite problem to a hollow one, so a fix aimed at reverb will not help it.
Why the room is so hard to take out
Here is the part that has frustrated people for as long as recording has existed. Reverb is not a separate sound sitting alongside your voice. Reverb is your voice, arriving again and again, out of time with itself.
That single fact defeats every conventional repair. You cannot cut a frequency band, because the reflections occupy exactly the frequencies your voice occupies. You cannot gate the gaps, because the tail is loudest exactly when the voice is loudest. Anything you subtract, you subtract from the voice too.
The standing answer in editing forums has been that this cannot be undone, and for the tools those answers were written about, it is correct. One frequently quoted verdict from an audio community moderator puts it plainly: you are trying to remove something that is identical to the signal you want, with no clean reference to work from. Elsewhere the same answer is given as an analogy about baked cakes.
De-reverb effects do exist, and they are genuinely useful, but they work by estimating how much of each frequency at each instant is reflection and pulling that portion down. Estimate low and the room stays. Estimate high and you take the body of the voice with it. That trade is where the thin, watery, hollowed out quality comes from, and it is why the honest advice for a badly reverberant take has always been to record it again.
What changed: rebuilding instead of subtracting
The newer approach does not try to subtract anything. It asks a different question: what was this voice before the room got to it?
A model trained on a very large amount of clean speech learns the shape of human speech, how a vowel decays, where a consonant stops, what a voice does between words. Given a reverberant recording it reads the words and the speaker out of it, then generates that speech again, clean. The reflections are not removed. They are simply never drawn the second time, because the output is not your file with parts taken away, it is new audio produced to match it.
This is why a room can come off at all. Nothing is being subtracted, so nothing has to be cleanly separable. The technical name is speech enhancement, and the generative kind is a genuinely different technique from the subtractive kind that shares the name.
It is also a real trade, and the next two sections are the honest version of it.
Does AI enhancement change your voice?
Slightly, and you should expect to hear it. The output is a rebuilt performance rather than your original waveform with parts removed, so it lands very close to the take without being identical to it. Same words, same speaker, same delivery. The texture is regenerated.
Whether that is acceptable depends on the job. For an interview, a podcast, a voiceover, a lecture, or anything where the point is to be understood, a rebuilt voice usually beats a voice with a room around it. For forensic work, or anywhere the recording is evidence and must remain unaltered, it is the wrong tool, and so is every other enhancement effect.
One practical consequence: the better your input, the less the model has to infer. A quiet room with a mediocre microphone rebuilds more faithfully than a loud room with a good one.
Why does cleaned-up audio sometimes sound robotic?
Two opposite failures share the word. A subtractive effect removed too much and left holes in the voice, which sounds watery, warbling or thin, and the fix is to apply less of it. A generative model could not read part of the input and rebuilt it too smoothly, and the fix there is a cleaner input rather than a gentler setting.
Neither is text to speech, which never had a recording underneath it at all. We take all of them apart, with a table for matching what you hear to the fault you actually have, in why AI-enhanced audio sounds robotic.
How to fix a recording that sounds like a room
We ship a free browser tool that does the rebuild: Enhance a Recording. Recordings up to 60 seconds, no account, nothing to install.
- Upload the audio or video file. WAV, MP3, M4A and FLAC all work, and so do MP4 and MOV: drop the clip straight from your camera or your timeline and your browser reads the audio track out of it, so only the audio is sent.
- Run it. There is nothing to tune. No threshold, no reduction amount, no noise profile to capture, because the model is not being aimed at the room.
- Compare. The result comes back beside the original at a matched level, so the difference you hear when you switch is the enhancement and not a change in volume.
- Take the track. One file, one voice, and the result link stops working after an hour.
Two limits worth knowing. It handles one voice at a time, so a second person in the room is not something it is built for. And it returns speech at 16 kHz, which can sound narrower than a full band original. The before and after are matched to the same bandwidth so the comparison stays honest.
What if there are two people in the recording?
Order matters. Enhance a conversation and you get two voices in a clean room, which is rarely what anyone wanted.
Keeping one person out of several is a different technique called target speaker extraction, and it runs first. Give it a few seconds of the voice you want, take the track it returns, then enhance that. Our free isolation tool does the first half, and the step-by-step guide covers choosing a sample. The field guide to diarization, separation and extraction maps how the techniques relate, and we go through what each editor effect actually does to two voices separately.
When to just record it again
Enhancement is a rescue, not a substitute for a good capture. Some cases are beyond it, and it is better to know before you upload.
- The room is louder than you are. If the reflections outweigh the direct sound, there may not be enough of the original left to read.
- Two people are talking at once. Isolate first, then enhance.
- The recording is evidence. A rebuilt voice is not an unaltered one.
- You can still record it again. Get closer to the microphone, put something soft in the room, and pick the smaller space. Thirty seconds of preparation beats any model, because reverb you never captured is the only reverb that comes off perfectly.
Why the room was ever in there
Every problem on this page comes from one root. Text arrives as characters and images as pixels, each of them editable one unit at a time. Sound arrives as a single undifferentiated wave, with every voice, every noise and every room fused into it. Nothing in the file marks where your voice ends and the tiled wall begins, which is why pulling one of them back out was called impossible for as long as it was.
Crele is working on exactly that: full control of speech and audio, on the recordings people actually have rather than the clean ones models get demonstrated on. Which voice is speaking, how that voice sounds, and whether it is who it claims to be. The free tools on this site are the part of it you can use today. More is coming soon. If you want early access to it, join the waitlist.