Why AI video for CTV fails at the delivery gate
Connected TV crossed a structural line this year, and AI video for CTV is now a live production question rather than a thought experiment. The IAB's 2026 Outlook Study, built on responses from more than 200 brand and agency buyers, projects connected TV ad spend growing 13.8% year over year while linear television declines 1.7%. Budget is moving to the living-room screen faster than production capability is, and the teams handed those briefs mostly spent three years optimising for a phone held at arm's length.
The failure mode is not what most people expect. Generated spots rarely get killed because a client hated the look. They get killed at ingest, by an automated conformance check that never rendered a single frame for a human. Variable frame rate from a generation tool, an audio track half a second shorter than the video, chroma subsampling a publisher will not accept, a peak that trips a loudness gate: none of these are creative problems, and all of them return the same rejection email.
That gap matters because the correction loop is expensive. A rejected CTV master usually means regenerating shots, not just re-exporting, because the fault is baked into what the model produced. Treating the delivery spec as the last step in the process is how a two-day turnaround becomes a two-week one.
The spec sheet is the brief, not the afterthought
Publisher specs for connected TV are unusually explicit, and they are worth reading before the first prompt rather than after the final export. Google Ad Manager's programmatic video spec is a good representative sample of what the broader ecosystem enforces: it generates MPEG4 or MOV with H.264 video and AAC audio, caps the long edge of the frame at 1920 pixels, and accepts uploads up to 1.9 GB or 20 minutes per file.
The constraints that trip generated work are further down the sheet. Frame rate must be constant, at 23.98, 24, 25, 29.97 or 30 fps matching the native rate. Chroma subsampling is 4:2:0 generally and 4:2:2 preferred. Audio must be two-channel at -23 integrated LUFS, PCM at 16 or 24 bit or AAC, minimum 192 Kbps at a 48 kHz sample rate, with audio duration matching the video exactly. Duration lands between 6 and 60 seconds, with 6, 15 and 30 the slots that clear the widest inventory.
Read that list as a production brief and several decisions make themselves. You are finishing in landscape, not reframing a vertical master. You are cutting to fixed durations, not to whatever the edit wants. You are committing to a frame rate before generation, because conforming variable-rate output afterwards costs quality. And you are budgeting an audio pass, because no generation tool hands you a compliant mix.

Generate at the fidelity floor, then route the job
A phone screen forgives a great deal that a 65-inch panel does not. The same clip that reads as clean at 6 inches shows compression mush in gradients, unstable edge detail and inconsistent grain once it is scaled to a wall and viewed from three metres. Upscaling a 720p generation to 1080p does not fix this; it redistributes the same information across more pixels and usually makes the softness more obvious.
The practical rule is to generate above your delivery resolution wherever the model allows it, and to reserve your highest-fidelity engine for the hero shots that hold on screen longest. Cheaper, faster engines are still fine for cutaways, product inserts and anything under a second. This is where choosing the right generation model for each job stops being a procurement debate and becomes a delivery constraint, because the model you pick determines whether you can hit 4:2:2 chroma and a 15 to 40 Mbps bit rate without visible banding.
Encode discipline matters just as much. Master to a high-bitrate intermediate, do your grade and conform there, and only then produce the H.264 deliverable. Exporting straight from a generation tool to a publisher-ready MP4 stacks two lossy passes on top of each other and throws away exactly the detail the big screen exposes.
Artifacts that survive a phone die on a big screen
Viewing distance changes which defects matter. Hand and finger errors, morphing background faces, drifting logos on props and text-like shapes that resolve into nonsense are all forgivable at thumbnail scale and glaring at broadcast scale. Temporal instability is worse still: a flicker in a two-second hold that nobody notices in a feed becomes the only thing anyone looks at in a living room.
The review process has to change with it. Check every cut on a television at realistic distance, not on a monitor at 40% zoom, and check it before you lock the edit rather than after the client has approved a preview link. Budget review time accordingly, and keep post-production fixes for common AI artifacts in the schedule rather than treating them as an exception, because on a 30-second CTV spot the odds that all shots come back clean are close to zero.
Pacing is part of the same problem. Rapid social-style cutting hides artifacts, which is why teams lean on it, but CTV convention favours two to four second holds and deliberate camera moves. A hold that long is a stress test. If a shot cannot survive three seconds without a viewer noticing something wrong, it is not a CTV shot.

Sound-on is the default and -23 LUFS is the number
Social video is designed for silence. Connected TV is not: the ad plays at whatever volume the household set for the programme, full screen, with nothing else competing. That inverts the sound priority, and it introduces a hard technical target that most performance teams have never had to hit.
EBU Recommendation R 128 defines the target programme loudness as -23 LUFS, and since version 3.0 the permitted deviation has been tightened to 0.5 LU where exact normalisation is not achievable. That is the same -23 integrated LUFS that publisher specs restate, so it is not a suggestion you can average your way around. Practically, it means a metered loudness pass on the finished mix, not a limiter slapped on the master bus, and it means a true-peak check so the mix does not clip on cheap television speakers.
Generated audio complicates this. Model output arrives at wildly inconsistent levels, and dialogue generated separately from the ambience will not sit together without a mix. Even with native audio generation improving, the mix still has to be conformed by a human before it clears ingest. Add the requirement that audio duration match video duration exactly and you have a step that belongs in the schedule, with a name and an owner.

Supers, legibility and the first two seconds
Nobody leans in to read a television. Legal lines, price points, offer terms and the brand lockup all have to survive being read across a room, and the temptation with generated footage is to let the model produce on-screen text, which it still does badly. Type belongs in the edit, composed against the plate, at sizes chosen for the viewing distance rather than for the canvas.
WCAG 2.2 gives the measurable floor. Success Criterion 1.4.3 requires a contrast ratio of at least 4.5:1 for normal text and 3:1 for large-scale text, where large-scale means at least 18 point, or 14 point bold, roughly 24 px and 18.5 px respectively. Those thresholds were written for screens at desk distance, so treat them as a minimum and scale up for the couch. Treating this as an encoding problem is the mistake; accessibility decisions taken at the creative stage are what keep supers readable across every seat in the room.
Placement follows the same logic. Keep supers well inside title-safe, avoid the lower band where some platforms overlay controls, and put the brand asset in early. Viewers cannot click away from a CTV pod, but they can look away, so establishing the brand within the first two seconds is standard practice rather than a nicety.
Run it as a gate, not a vibe check
Turn all of the above into a pass or fail list that runs before anything leaves the building. Container and codec correct. Resolution 1920 x 1080 at 16:9. Frame rate constant and matching the native rate. Bit rate inside the recommended band. Chroma subsampling at 4:2:0 or better. Integrated loudness at -23 LUFS with a clean true peak. Audio duration equal to video duration. Duration on a standard slot. File size under the ceiling. Supers inside title-safe and above the contrast floor. Every shot reviewed on a television.
Automate the measurable half. A short ffprobe and loudness script will catch frame rate, chroma, bit rate, duration mismatch and integrated loudness in seconds, which leaves human attention for the judgement calls: whether a hold survives three seconds, whether the mix sits, whether the brand reads. Bolt these onto the five-gate QC checklist you already run rather than inventing a parallel process, because a second checklist nobody owns is a checklist nobody runs.
The payoff is that versioning becomes cheap. Once one master conforms, cutting audience-specific variants is a re-render against a known-good template rather than a fresh negotiation with a publisher's ingest system. That is the actual argument for generative production on connected TV: not that a spot costs less to make, but that the second, fifth and twentieth version cost almost nothing once the delivery gate is solved.
Put the framework into production
These related pages connect the article’s planning advice to a specific commercial scope.
References
- IAB 2026 Outlook Study Forecasts 9.5% Growth in U.S. Ad SpendInteractive Advertising Bureau
IAB's 2026 Outlook Study, based on more than 200 brand and agency buyers, forecasts connected TV ad spend growth of 13.8% year over year in 2026 while linear TV declines 1.7%.
- Video creative specifications in Google Ad ManagerGoogle Ad Manager Help
Google Ad Manager's video creative specification generates MPEG4 or MOV files with H.264 video and AAC audio, caps the long edge of the video at 1920 pixels, normalises audio to a target of -24 LKFS per ATSC A/85, and accepts uploads up to 1.9 GB or 20 minutes per file.
- EBU Recommendation R 128: Loudness normalisation and permitted maximum level of audio signalsEuropean Broadcasting Union
EBU R 128 sets the target programme loudness level at -23 LUFS, with the permitted deviation tightened to 0.5 LU from version 3.0 onwards where exact normalisation is not achievable.
- Understanding SC 1.4.3: Contrast (Minimum)W3C Web Accessibility Initiative
WCAG 2.2 Success Criterion 1.4.3 requires a contrast ratio of at least 4.5:1 for normal text and 3:1 for large-scale text, defined as at least 18 point or 14 point bold, approximately 24 px and 18.5 px.
