Skip to content
Working with the tools, not about them

Images & Audio

Sound that sits under a voice has to be built to lose the argument

Generated music and effects are usable as background, and the properties that decide whether they work are structural rather than musical.

By Aarav Sinha4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Background audio has a job and it is not to be interesting

Music placed under narration exists to fill silence, mark a transition and set a register. It succeeds by being unnoticed. That makes it almost the opposite of what a request for good music produces, which is something with a melody, a development and a moment where it goes somewhere.

A melodic line competes with speech directly, because both occupy the same range and both invite the listener to follow a sequence. The result is a listener who is doing two things at once and retaining neither, and the usual diagnosis is that the music is too loud when the problem is that it is too eventful.

So the request should ask for what a bed actually is: consistent texture, no strong melody, no dramatic change in intensity, nothing that resolves. Those are unglamorous words and they describe the thing that works.

Length and joins are the practical problem

Generated audio arrives in fixed lengths that won’t match your narration, and this is where most of the work goes. A thirty-second piece under a four-minute segment has to loop, and a loop is audible the moment the join doesn’t match in level, in reverberation tail or in where the beat falls.

Two things help. Ask for material that is deliberately static, because static material loops without a seam. And do the looping in an editor with a crossfade, rather than expecting a clean cut, since a short overlap hides an enormous amount.

The alternative is to generate several distinct pieces and let each section of the programme have its own, with the changes placed at points where the narration also changes. A change of music at a transition sounds intentional. The same change in the middle of a paragraph sounds like a mistake, because it is one.

Effects need a place, not just a name

Individual sounds — a door, rain, a room tone, traffic — are among the more reliable things to generate, particularly when they will be heard for a second under something else. They are short, and short means fewer opportunities for the failure to accumulate.

What they need is spatial description rather than identification. Rain on a window from inside a room is a different sound from rain outdoors, and both are different from rain on a metal roof. Distance, surface, enclosure and whether the listener is inside or outside carry the realism.

Room tone deserves particular mention because it is the least obvious and most useful. A recording with genuine silence between sentences sounds wrong; a low continuous bed of room noise makes cuts disappear. It is also the easiest kind of audio to generate acceptably, since nobody is listening to it.

The check is always the same and it isn’t technical. Play it under the voice, at the level it will actually sit, and listen to whether you can still follow the words. Effects judged in isolation are judged in a situation that will never occur.

Mixing decides more than generation does

The difference between amateur and competent background audio is almost entirely level, and level is not a generation parameter. Under speech, music generally has to sit far lower than instinct suggests, and it has to duck further whenever a sentence begins.

Doing this properly is a job for an audio editor, and it is a shallow skill with a large payoff: set the speech level first, bring the bed up until it is just audible, then reduce it. Anything that makes you strain to hear a word is too loud regardless of what the meter says.

Frequency matters too, though it can be described simply. Beds that are dense in the same range as a voice will muddy it even at a low level, which is why a sparse, low or airy texture works better under speech than a full arrangement, whatever its quality.

Where this stops being the right approach

Music that has to carry a scene on its own — a theme, something with a hook, anything a listener is meant to remember — is a different requirement, and generated material rarely satisfies it. It tends to be pleasant, competent and forgettable, which is fine for a bed and fatal for a theme.

There are also questions of rights and disclosure that vary by platform and by jurisdiction, and they are worth resolving before a piece is published rather than after. Where a recording will be distributed commercially, the terms attached to the material matter as much as how it sounds, and that is a question for whoever handles the licensing.

The honest position is that this is a good way to obtain the audio nobody notices, and a poor way to obtain the audio somebody remembers. Most projects need far more of the first than the second.

Common questions

Why does generated music fight with narration?

Because a request for good music produces melody and development, and both compete with speech for the listener’s attention. A bed should be static, unresolved and texturally consistent, which is what the request has to ask for explicitly.

How do I loop a short piece cleanly?

Generate deliberately static material, then loop it in an editor with a short crossfade rather than a hard cut. Alternatively, give each section of the programme its own piece and place the changes where the narration also changes.

What level should background audio sit at?

Lower than instinct suggests, and it should drop further when a sentence starts. Set the speech first, raise the bed until it is just audible, then reduce it. Judge it playing under the voice, never in isolation.

Images & Audioaudiomusicsound designediting
Aarav Sinha
Features writer, Prompt After Prompt

Aarav covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.