What AI Video Sound Design Actually Covers
AI video sound design is the pass that turns a generated clip into footage people believe. Picture models now ship motion and lighting that survive a phone screen; the audio track rarely survives a first listen. Build the sound as a separate deliverable — ambience, effects, loudness and captions — or the cut reads as a render.
Generated footage arrives silent. Not quiet — silent, in the sense that there is no production sound to fall back on. A camera shoot captures room tone, fabric, footsteps and the small imperfections that tell a viewer where a scene physically is. A generation returns pixels. Every layer that normally arrives for free has to be built, and the layers are not equally visible once they are missing.
The first layer is the ambience bed. Run one continuous bed under every shot in a sequence, chosen for the location rather than the moment: traffic outside an office, refrigeration hum in a kitchen, wind across an exterior. Because the shots were generated separately, a shared bed is the cheapest way to make six unrelated clips read as one place. Spot effects come second — a footstep matched to the gait, a hand setting down a glass, cloth moving with a turn.
Third is perspective. A sound designed for a close-up and dropped onto a wide shot flattens the frame, because the reverb and distance cues still say close. Fourth is balance: voice against music against effects, with the voice usually winning. Teams that skip the first three and start with music end up with a cut that looks finished and sounds like a demo reel, and the fix is the same discipline that governs {{link}}.
A useful test is to strip the music from a finished cut and listen to what is left. If the answer is almost nothing, the sequence is being held together by a library track rather than by a designed space, and it will not survive either a platform transcode or a mute button.
Teams that skip the first three and start with music end up with a cut that looks finished and sounds like a demo reel, and the fix is the same discipline that governs post-production repair for AI video.

Generative Audio Does Not Remove the Pass
The model layer now ships picture and sound together, and that is genuinely useful — it removes the blank timeline problem. It does not remove the pass. Generated audio arrives with no ambience bed running underneath it, no loudness target, and no guarantee that the effect lands on the frame the edit needs. It is closer to a scratch track than to a mix.
Treat generated audio as source material and route it through the same decisions as recorded sound: {{link}} is a capability question, not a finishing answer — the mix still has to be assembled, levelled and checked. Keep a stem of everything the model produced so the bed and the spots can be rebalanced later without regenerating the picture.
There is a provenance argument too. Sound that a model produced carries the same questions as the picture: which system generated it, under what terms, and whether the output can be reused in a later cut. Logging the audio generation alongside the video generation costs nothing at the time and is close to impossible to reconstruct six months later.
Lip sync is the sharpest version of this. A generated line that lands two frames late does more damage than no line at all, because the viewer reads the mismatch as a fake person rather than as bad audio. Where the line carries the claim, record it separately and time the cut to the recording instead of the other way round.
Treat generated audio as source material and route it through the same decisions as recorded sound: simultaneous audio-visual generation is a capability question, not a finishing answer — the mix still has to be assembled, levelled and checked.
Loudness Is a Delivery Contract, Not a Preference
EBU R 128 normalises programme loudness to a target level of -23.0 LUFS, measured over the programme in its entirety rather than with emphasis on foreground elements such as speech, music or sound effects. The tolerance has been ±0.5 LU since version 3 of the recommendation, and ±1.0 LU is permitted only where reaching the target is not practically achievable. During production the true peak level must not exceed -1 dBTP, with a measurement tolerance of ±0.3 dB.
That bites harder on AI video than on a conventional cut. A sequence assembled from separately generated shots, each carrying its own generated audio, has no shared level and no shared acoustic space. Normalising clip by clip will land the number on every file and still leave the programme wrong, because the standard measures the whole programme rather than its parts.
The practical fix is to make loudness a gate rather than a preference: {{link}} — one measurement, on the finished file, before anything ships. Measure after the last render, not at the rough cut, because any re-export that changes the balance changes the number.
The practical fix is to make loudness a gate rather than a preference: a QC gate for AI video — one measurement, on the finished file, before anything ships.

The File Has to Survive the Platform
YouTube's recommended upload encoding settings are unusually explicit about audio: an MP4 container with the moov atom at the front, an audio codec of AAC-LC or Opus or Eclipsa Audio, a sample rate of 48 kHz, and channels of stereo, stereo plus 5.1, or stereo plus Eclipsa Audio. Recommended audio bitrates are 128 kbps for mono, 384 kbps for stereo, 512 kbps for 5.1, and 128 kbps per channel for Eclipsa Audio.
Export at 48 kHz and let the platform decide what to transcode; do not hand it a file that an editing tool has already resampled. Immersive audio is a real delivery option now rather than a novelty, but it carries its own arithmetic — per channel, not per file — so it belongs in the quote rather than as a surprise at export. The same discipline applies to frame rate: content should be encoded and uploaded at the frame rate it was recorded, and any interlaced material should be deinterlaced before upload, because a generated clip that an editor re-times will drift against its own audio.
Sound is also the layer that travels worst across markets: {{link}} — a separate audio track is a deliverable in its own right, and the music and effects stem should be exported alongside it so the next market is a mix, not a regeneration.
Sound is also the layer that travels worst across markets: AI video localization — a separate audio track is a deliverable in its own right, and the music and effects stem should be exported alongside it so the next market is a mix, not a regeneration.
Assume Nobody Hears It
WCAG Success Criterion 1.2.2 requires captions for all prerecorded audio content in synchronized media, and it is explicit about what those captions carry: not only dialogue, but who is speaking and non-speech information conveyed through sound, including meaningful sound effects. Every decision made in the sound pass therefore has a text twin that someone has to write down. In practice that means the sound designer and the caption writer are working from the same cue sheet, and that the cue sheet is written while the mix is being built rather than after it.
That is not only an accessibility obligation. Consumer research published in 2026 found that when an ad cannot be skipped, 44% of viewers keep watching, 34% turn away and 17% mute it. A cut whose storytelling depends entirely on a sound bed loses that share outright, which makes the sound design and the caption track the same job viewed from two ends.
The working rule is simple: watch the final cut once with the sound off. If the beat still reads, the sound pass is doing its job of adding belief. If it does not, the problem is usually in the picture rather than in the audio.

A Sound Pass Checklist for AI Video
Build the ambience bed before the spot effects, and run one bed across the whole sequence rather than one per shot. Cut on motion and hold on emotion; most generated sequences are cut too fast, which makes them read as a reel instead of a scene. Keep generated dialogue out of the final unless its timing has been checked against the cut frame by frame.
Normalise the finished programme once, at -23.0 LUFS, and check the true peak at -1 dBTP. Export at 48 kHz in the codec the destination asks for, with the bitrate budget agreed before the job is quoted. Write the caption track from the sound design, including the effects that carry meaning. And keep a record of which model produced which piece of audio, because six months later that is the only way to answer where a sound came from. None of this is expensive. It is an afternoon of work on a thirty-second cut, and it is the difference between a clip that gets watched and one that gets scrolled past.
Put the framework into production
These related pages connect the article’s planning advice to a specific commercial scope.
References
- EBU R 128 (2023, v5.0) - Loudness normalisation and permitted maximum level of audio signalsEBU
Programme Loudness is normalised to a Target Level of -23.0 LUFS with a tolerance of ±0.5 LU since version 3, and the True Peak Level of a programme shall not exceed -1 dBTP during production.
- Recommended upload encoding settings - YouTube HelpGoogle
YouTube recommends MP4 with audio codec AAC-LC or Opus or Eclipsa Audio, 48 kHz sample rate, and audio bitrates of 128 kbps mono, 384 kbps stereo and 512 kbps for 5.1.
- Understanding Success Criterion 1.2.2: Captions (Prerecorded) - WCAG 2.2W3C
Captions are required for all prerecorded audio content in synchronized media and must identify who is speaking and include non-speech information conveyed through sound, including meaningful sound effects.
- 2026 US Media Consumption Report - The Attention EconomyAttest
When an ad cannot be skipped, 44% of viewers keep watching, 34% turn away and 17% mute it.
