Why AI Video Accessibility Fails at the Generation Stage

AI video accessibility is not a captioning task you bolt on after picture lock. It is a set of decisions about what the audio track carries, what the picture carries alone, and what happens to the message when a viewer cannot use one of those two channels. Generative pipelines make those decisions earlier and far more often than traditional production ever did, usually without anyone noticing that a decision was made at all.

The failure pattern is consistent. A model produces a visually dense spot with no dialogue, so the offer ends up living entirely in on-screen supers. The supers are generated too, so the typography is approximate and the contrast is whatever the model felt like that run. There is no script document, because the prompt was the script. By delivery, three separate access requirements have quietly become impossible to satisfy without regenerating shots.

None of this is exotic. It is the predictable result of compressing preproduction into a prompt. The fix is equally unglamorous: name the access layer in the brief, keep a written source of truth alongside every generation run, and treat two specific artifacts, garbled supers and frame-level flicker, as delivery blockers rather than polish items.

Write the Audio Description Into the Script, Not After the Edit

In June 2026 the UK advertising regulator published the clearest statement yet on where accessibility belongs. CAP and BCAP reported that the strongest theme from their 2025 stakeholder engagement was accessibility by design: provisions work better when they are considered at the beginning of the creative process rather than retrofitted once an ad has already been created.

The companion rule has teeth. Where an ad uses accessibility provisions, any material information, meaning the information an average consumer needs to make an informed transactional decision, has to reach the audience through those provisions as well. A headline price qualified by on-screen text about contract length must appear in the audio description too, or the ad misleads a visually impaired viewer.

That decision belongs upstream. A creative brief for AI video already forces you to name the objective, the audience, the mandatories and the claims that need qualifying, so adding one line for the access layer costs almost nothing and saves a re-record later.

The practical version for generative work is a binary choice made at brief stage: audio-led or visually led. Audio-led means narration carries the story and the qualifications, so description is largely built in. Visually led means you are committing to a separate description track and must leave gaps in the mix for it, a constraint the prompt needs to know about before a single frame renders.

Diagram of a script page branching into an audio-led waveform path and a visually led filmstrip path with an empty description track

Captions Need a Source of Truth the Model Cannot Provide

Models that generate native dialogue create an odd gap. There is speech in the deliverable but no script behind it: whatever the character says was invented during sampling, and the only record is the rendered audio. Captioning then degrades into transcribing your own output, which is slower, less accurate and impossible to version cleanly.

Keep a written line list per shot and treat it as the caption source. Write the intended dialogue, generate against it, then reconcile whatever the model actually produced back into the list before the cut leaves the edit. That same list doubles as your dub script, your claims record, and the text an audio description writer works from.

It is also the spine of any multi-market rollout, and a localization playbook that starts from a locked master line list will survive far more versions than one that starts from a rendered mix.

Do not treat platform auto-captions as a deliverable. YouTube lets you upload a caption file, auto-sync from a transcript, or type captions manually, and its automatic captions are generated in the video's default language only. For a paid placement carrying material claims, the reviewed file is the deliverable and the automatic track is a fallback nobody has signed off.

On-Screen Text Is the Weakest Link in Generated Footage

Generated typography is still unreliable, and the failure is not merely aesthetic. If a super carries a legal qualification and the model renders it with a dropped character or a drifting baseline, the qualification is arguably not communicated at all, and no caption track can repair text that lives inside the picture.

Composite supers instead of generating them. Render the plate clean, then bring legal lines, price points and end-frame lockups in as a graphics layer you control, with a known typeface, a measured contrast ratio and a defined safe area. That is the only way to guarantee the qualification survives a re-grade, a reframe and a platform re-encode.

It solves the reframing problem too. A 16:9 master reframed to 9:16 will crop burned-in supers in ways nobody previews, while a graphics layer can be repositioned per aspect ratio without touching the generated plate. Treat any burned-in text inside a generated shot as a defect to be replaced, not an asset to be preserved.

Side by side comparison of a smeared generated caption block versus a clean separate graphics layer with safe area guides

Treat Flicker as an Accessibility Defect, Not a Polish Item

Temporal instability is the signature artifact of diffusion video: brightness pumps between frames, textures boil, a practical light strobes across a cut. Teams file this under quality. At sufficient amplitude it is also a safety and accessibility issue, and unlike most craft judgements it comes with a published numeric threshold.

WCAG 2.2 Success Criterion 2.3.1 requires that content not contain anything that flashes more than three times in any one second period, unless the flash stays below the general flash and red flash thresholds. A general flash is defined as a pair of opposing changes in relative luminance of ten percent or more of maximum, where the darker state sits below 0.80 relative luminance.

Most AI flicker is low amplitude and passes comfortably. The cases that do not are strobing practicals, rapid cut sequences assembled from unstable generations, and saturated red transitions, which the criterion treats separately and more strictly. Measure rather than eyeball it: a luminance scope parked over the questionable range answers the question in seconds.

Where a shot does fail, this is a stabilization and deflicker problem before it is a regeneration problem, and the usual post-production fixes for AI artifacts apply in exactly the same order.

Schematic of film frames above a luminance graph where three pulses cross a marked threshold line within one interval

Where Access Requirements Bind and Where They Are Only Expected

The regulatory picture is uneven, and knowing which side of the line a campaign sits on changes the budget conversation. In the EU, the European Accessibility Act covers a defined list of products and services that includes access to audio-visual media services and related consumer equipment, which pulls distribution and platform layers into scope even where an individual spot is not named directly.

In the UK there is currently no legislative requirement for accessible advertising, and the regulator has said plainly that it is not in a position to mandate one. It still enforces that material information must reach viewers through whatever access provisions an ad does carry. Voluntary in principle, enforceable the moment you opt in.

Practically that means two budget lines rather than one. The first is what a market legally compels. The second is what a client's own inclusion commitments and platform contracts compel, which on most large accounts is the binding constraint long before legislation is. Scope both at brief stage, because discovering either at delivery is how a spot slips a week.

Run the Access Pass While the Cut Is Still Cheap to Change

Fold access checks into the review you already run rather than inventing a parallel process. Five checks cover most of the exposure: caption file present and reconciled against the line list, audio description written or the narrative confirmed audio-led, all material information present in both channels, supers composited with measured contrast, and flicker measured on any shot with visible pumping.

Every one of those belongs in the same review where you already validate continuity and provenance, so extend your pre-delivery QC checklist rather than bolting on a second sign-off nobody actually runs.

Assign owners explicitly. The creative lead owns audio-led versus visually led. The producer owns whether description is budgeted. The editor owns the caption file and the supers layer. Finishing owns flicker measurement. Unowned checks are precisely the ones that get skipped at six o'clock on a delivery day.

Running this pass at brief and again at picture lock costs perhaps an hour. Discovering at delivery that the offer exists only inside a garbled generated super, with no line list to caption from and no gap in the mix for description, costs a reshoot you cannot shoot, only regenerate, and hope the model cooperates.

Put the framework into production

These related pages connect the article’s planning advice to a specific commercial scope.

Short-form ad productionTurn hook strategy into platform-ready creative variants.AI UGC productionBuild creator-style openings into a controlled testing system.

References

  1. Accessibility provisions and material information in advertisingCommittee of Advertising Practice (CAP/BCAP)

    CAP and BCAP report that the strongest theme from their 2025 stakeholder engagement was accessibility by design, and that where an ad uses accessibility provisions, material information must be communicated through those provisions as well.

  2. Understanding SC 2.3.1: Three Flashes or Below ThresholdW3C Web Accessibility Initiative

    WCAG 2.2 Success Criterion 2.3.1 requires that content not contain anything that flashes more than three times in any one second period unless the flash is below the general flash and red flash thresholds.

  3. European Accessibility Act (EAA)European Commission

    The European Accessibility Act covers a defined list of products and services that includes access to audio-visual media services such as television broadcast and related consumer equipment.

  4. Add subtitles & captionsYouTube Help

    YouTube supports uploading a caption file, auto-syncing from a transcript, or typing captions manually, and its automatic captions are produced in the video's default language only.

Related reading

How to Write a Creative Brief for AI Video That Actually DeliversThe AI Video Localization Playbook: One Master, Twelve MarketsFixing AI Video Artifacts in Post: A Production PlaybookThe AI Video QC Checklist: Five Gates Before a Cut Ships