Blog · 8 min read

How to fix a mistake in a recording without re-recording it

The mistake is always found after everyone has gone home. Five ways a wrong word gets fixed today, what each one needs from you, and the one that no longer needs a studio.

By · Published Sep 16, 2026

The mistake is always found later. The interview is over, the studio is packed up, the guest is on a plane, and somewhere in the middle of a good take is a wrong name, a wrong number, or a sentence that came out backwards.

You have more options than re-recording it, and they are not equally good. Here is what people actually do, what each one costs, and what changed in the last couple of years.

Cut it out, if the words can go

The free fix, and the one to try first. Find the mistake, select it, delete it, and close the gap.

It works better than people expect, as long as you cut in the right place. Cut on a breath or a pause, never inside a word. The start and end of a breath are the quietest moments in a sentence, so a join there has the least to give away. Add a short crossfade over the join, a few milliseconds, and most editors will do it for you automatically.

Where it stops working is when the words are load-bearing. You can lose a stumble, a false start, a repeated word or a tangent. You cannot lose a price, a name or a “not”, because the sentence still needs them. Cutting also changes the rhythm of a sentence, and a line that loses its middle can end up sounding hurried even when the join itself is clean.

Fill the gap with room tone

Here is the part that surprises people the first time: an edit can be audible even when nothing is wrong with the words.

Every recording has a quiet layer under the speech. Air conditioning, traffic outside, the hiss of the microphone itself, the sound of the room being a room. Film and dialogue editors call it room tone, and on a professional shoot somebody records 30 to 60 seconds of it with the same microphone in the same position, with everyone standing still and saying nothing.

The reason is that a cut can drop that layer to nothing for a moment. The words either side are fine, the background disappears and returns, and the listener hears the edit without being able to say what they heard. Filling the gap with room tone puts the layer back, and the join vanishes.

If you did not record any, you can usually steal a second or two from a pause elsewhere in the same recording and loop it. If you have nothing, this is why your cuts breathe.

Room tone is also worth knowing about for a reason that arrives later on this page: silence is not silence, and anything dropped into a recording has to bring the right background with it.

Punch and roll, if you are still in the chair

The best fix is the one made before you leave. Audiobook narrators and voiceover artists use a technique called punch and roll, and it is the reason a good narrator does very little editing.

You stop as soon as you flub. You roll back to the last clean phrase, play it into your headphones, and start speaking again in time with it, recording over the mistake as you go. The correction happens in the same room, on the same microphone, at the same distance, in the same voice, minutes after the words around it.

There is nothing to match, because nothing has changed yet. That is the whole trick, and it is also the limit: it only works while the session is still running. Once the microphone is packed away, punch and roll is not available to you any more, and every remaining option on this page is some attempt to recreate what you had.

Record a pickup line later

The obvious move once the session is over: set the microphone back up and say the line again.

It can work, and it often does not, because almost nothing about the second recording is the same as the first. A different room, or the same room with the window open. The microphone a few inches further away, or at a different angle. A voice that woke up earlier, or later, or has a cold. And a line delivered on its own, cold, rather than arriving in the middle of a thought.

Any one of those is enough to hear. Together they are why a pickup so often sounds like what it is.

If you have to record one, get as close to the original setup as you can: same place, same microphone, same distance, same time of day if the room changes through it. Then match the level, and put the pickup somewhere the sentence naturally breathes rather than in the middle of a run.

ADR, the expensive version of the same idea

Film and television have an industrial-grade version of the pickup, called automated dialogue replacement. When the location sound is unusable, the performer is brought into a studio weeks or months later and re-records the lines while watching the scene, in time with their own lips.

It is expensive, it needs the performer available, and the performance is the hard part: delivering a line the same way months after the moment, without the scene around you, is genuinely difficult work.

But the detail worth taking away is technical. A studio has no room sound of its own. That is what makes it a good place to record. So once the clean line is captured, the location’s acoustics have to be added back on top to make it sit in the scene. The industry’s most expensive fix for a wrong line ends with somebody carefully painting the original room back onto it.

What every one of these has in common

Look at them together and the pattern is hard to miss.

FixWhat it costsWhat it needs
Cut it outNothingThe words to be expendable, and a pause to cut on
Room tone fillNothing, if you have the toneSomebody to have recorded the room on the day
Punch and rollNothingTo still be in the session
Pickup lineAn hour and a setupThe same room, microphone and voice, later
ADRA studio and a performerBudget, and the room added back afterward

Every one of them is an attempt to reproduce the conditions of the original take. And every one of them, except cutting, has to be arranged before you need it: the room tone recorded on the day, the session still running, the performer still available, the budget approved.

The mistake, meanwhile, is usually found long after all of that is gone.

The new approach: regenerate the words

The last few years added an option that does not depend on any of that. Edit the transcript, and a model regenerates the words in the speaker’s voice. Text-based editors have shipped versions of this for a while now.

It is a real change. No studio, no performer, no session, no foresight. The voice match, which sounds like it should be the hard part, is usually good.

What people report is that the edit still shows, and the reasons line up exactly with everything above. The regenerated words arrive clean. They were made somewhere with no room in it, no microphone distance and nothing underneath, so they drop into a recording that has all three. For half a second the listener is somewhere quieter and closer, then back again. It is the room tone problem wearing different clothes, and the ear catches a change of room faster than a change of voice.

On a rough take there is a second failure. When the model has very little clean voice to work from it can come back warbling or unnaturally smooth, or substitute a word that fits the sound it heard rather than the one you typed. That family of faults is worth knowing in its own right, and we went through it in why AI-enhanced audio sounds robotic.

So the generative approach solved the part everyone expected to be hard, the voice, and left the part the rest of this page is about: matching the recording.

What a seamless edit actually requires

Put the two halves together and the specification writes itself. New words have to arrive in the conditions of the take, not the conditions of a studio:

  • The room. The same reflections, the same sense of the space.
  • The distance. A word that is suddenly closer to the microphone is as obvious as one that is suddenly louder.
  • The background, running through the join. Not added afterward. Continuous from the words before to the words after, the way room tone is.
  • The delivery of the sentence. A word said on its own has the wrong stress and the wrong length for the middle of a phrase.
  • Joins where the sentence breathes. The same rule as cutting, for the same reason.

That is a different job from cloning a voice, and it is the one we are building. Speech editing is coming soon: highlight the words that are wrong, type what should have been said, and get the recording back with no audible seam, on the rough recordings people actually have rather than studio takes. If you can hear where the edit is, it is not finished.

Questions

How do you fix a mistake in a recording without re-recording it?
If the words can go, cut them out at a breath and let the background run across the join. If they cannot, your options are a pickup line recorded in the same room, or a tool that regenerates the words in the speaker's voice. Cutting is free and works today. Everything else is a matter of how well the fix matches the conditions of the original take.
What is punch and roll?
A way of fixing a flub while you are still recording. You stop, roll back to the last clean phrase, listen to it in your headphones, and start speaking again in time with it, recording over the mistake. Audiobook narrators use it because the fix is made in the same room, on the same microphone, in the same voice, so there is nothing to match afterward. It only works before you leave the chair.
What is room tone and why do editors record it?
Room tone is a stretch of a location recorded with nobody talking, usually 30 to 60 seconds, captured with the same microphone in the same position as the dialogue. Editors use it to fill gaps, so a cut does not drop to pure silence. Every recording has a quiet layer under the speech, and when that layer stops for a moment the listener hears the edit even though nothing else changed.
What is ADR?
Automated dialogue replacement: bringing a performer into a studio after the shoot to re-record lines and syncing them to the picture. Film uses it when the location audio is unusable. It is expensive, it needs the performer back, and because a studio has no room sound of its own, the location's acoustics have to be added back on afterward to make the new line match the scene.
Why does a pickup line sound different from the rest of the recording?
Because almost nothing about it is the same. A different room, a microphone at a different distance, a voice on a different day, and a line delivered on its own rather than in the middle of a sentence. Any one of those is audible. The closer you can get to the original setup, the same place, the same microphone, the same distance, the better a pickup sits.
Can AI replace a word in a recording?
Yes. Text-based editors can regenerate a word in the speaker's voice from the transcript, and the voice match is usually good. What people report is that the edit still shows: the regenerated word arrives clean while the recording around it is not, and on a rough take it can warble or come back as the wrong word. The voice is the solved part. The recording around it is not.
Why does the regenerated word sound cleaner than the rest of the take?
Because it was made somewhere with no room in it. Your recording carries a room, a microphone distance and a layer of noise under every word. A generated word carries none of that, so for half a second the listener is somewhere quieter and closer, then back again. It is the same problem room tone solves for cuts, and the ear catches a change of room faster than a change of voice.
Can you change what someone says in a video?
The audio, yes, the same way as any recording. The lips are a separate problem: in a close-up the mouth will still be saying the old words. Voiceover, narration and shots where the mouth is not the subject are fine. A talking-head close-up needs the picture changed too, which is a different kind of tool.

What to use today

Until it ships, the honest summary: cut if the words can go, fill the gap with room tone, punch and roll while the session is live, and record a pickup in the same room if you truly need the words replaced.

Crele builds Audio AI for the recordings people actually have. Two pieces of the engine are free to use in the browser right now, with no account: Isolate a Speaker keeps one voice and removes everyone else, useful when the problem is not what was said but who else was talking, and Enhance a Recording takes a rough recording of one person and returns it clean. Speech editing is the next piece. To hear it first, join the waitlist on its page.

Upload a recording now

Sign up for free and experience both models on your own recordings with longer files and downloads