Repairing a recording is a steadier use than synthesising one

The reliable audio work is corrective
Attention goes to generated voices, and the more dependable gains sit in processing recordings that already exist. Removing background noise, evening out levels, reducing a room’s echo, separating speech from music: each of these has a clear objective and a result you can judge in seconds by listening.
That is the important structural difference. A cleaned recording can be compared against the original, so a bad outcome is obvious immediately. Synthesised speech has nothing to compare against, which means errors survive until a listener notices them.
The gains are also large in ordinary terms. Interviews recorded in unsuitable rooms, video calls with one participant on a poor connection, footage where a fan was running — material that would once have been unusable is now frequently recoverable, and recovering it takes minutes rather than an afternoon.
Aggressive cleaning has a characteristic sound
Every noise reduction is a trade. Push it far enough and the noise disappears along with the top of the voice, leaving speech that sounds thin, slightly underwater and oddly gated between words. Listeners cannot name the artefact but they register the recording as cheap.
The practical rule is to apply less than the maximum and to listen at the setting where noise is still faintly present. A little residual room tone is normal and reads as natural; complete silence between phrases does not, because real rooms are never silent.
Check on the equipment your audience will use. Artefacts that are inaudible on good headphones can be very obvious through a phone speaker, and the reverse happens too. Judging a cleanup on one playback path is how a fix that seemed complete turns out to have introduced a new problem.
Transcription changed what is possible, within limits
Automatic transcription is now accurate enough that editing audio by editing text is a normal way to work, and searching a year of recordings for a phrase takes seconds. For anyone producing spoken material regularly, this is a bigger practical change than voice generation.
The accuracy is uneven in predictable ways. Names, technical terms, product names and numbers are the weak points, along with strongly accented speech, overlapping speakers and poor recordings. Those are precisely the parts most likely to matter, which makes a targeted check worthwhile even when the overall transcript looks clean.
Supplying a short list of expected names and terms in advance helps considerably where the tool accepts one. It is a small preparation step that removes the most common corrections, and it costs less than fixing the same misspelled surname forty times.
Text-based editing hides the joins it makes
Cutting a sentence by deleting it from a transcript is efficient and occasionally deceptive. The audio is joined at the deletion, and if the surrounding breath and room tone do not match, the edit is audible. Reviewing the joins by ear rather than trusting the text is the necessary habit.
There is an editorial dimension too. Removing filler words and pauses makes a speaker sound more articulate than they were, which is normal practice in edited interviews and becomes something else when the removed material changes what was meant. A hesitation before an answer can be information.
The line most people accept is that cutting for length and clarity is fine and reordering to change meaning is not. Where the recording is an interview with someone, the plainest safeguard is to keep the original and to be willing to show it.
Say what is synthetic, and treat voices as personal
Where generated speech is genuinely useful — a routine notice, a short label, a draft narration for timing before a real recording — the work is bounded and the stakes are low. Those are good uses and there is no reason to avoid them.
A person’s voice is a different matter from a generic one. Recreating someone’s speech without their agreement is a serious act regardless of how easy it has become, it is treated differently in different jurisdictions, and consent should be specific about what the voice may be used for rather than general.
Disclosure conventions are still forming, but the direction of travel is towards labelling, and being early rather than late costs nothing. If a listener would be surprised to learn that a voice was synthetic, that surprise is the signal that it should have been said in advance.
Common questions
How much noise reduction is too much?
Enough that the voice loses its top end or that gaps between words become completely silent. Real rooms have a faint noise floor, so leaving a little audible is more natural than removing everything, and checking on a phone speaker as well as headphones catches artefacts you would otherwise miss.
Where do transcripts usually go wrong?
Names, technical terms, product names and figures, plus overlapping speech and poor recordings. Those are also the parts most likely to matter, so check them specifically rather than skimming the whole transcript for general accuracy.
Is editing an interview by cutting the transcript acceptable?
Cutting for length and clarity is standard practice; reordering or splicing in a way that changes what someone meant is not. Listen to every join, since a text edit can leave an audible seam, and keep the original recording so the edit can be shown to be fair.
Reporter, Prompt After Prompt
Devika has been reporting on prompt craft, writing with ai, images & audio since long before it was fashionable and reads the small print so you do not have to.