Why AI video on-screen text comes out garbled

AI video on-screen text still fails, and the reason is structural: generative models paint letterforms as pixels instead of setting type from a font. The practical fix is not a better prompt. Keep words out of the generated frame entirely, then add every super, caption and legal line as a real type layer in post.

A text-to-video model has no font engine. It was never told that a lowercase letter has a fixed skeleton; it has only seen millions of frames in which letter-shaped textures appear, so it reproduces the texture and guesses the structure. Latent compression makes it worse, because most video models generate inside a compressed feature space, and the thin strokes and tight counters that make a glyph readable are exactly the detail that gets smeared on the way back to pixels.

Motion raises the difficulty another level. A still image only has to get the word right once. A clip has to hold the same word stable across every frame while the camera moves and the light changes, and per-frame jitter that would pass unnoticed on a brick wall becomes visible flicker in a wordmark. Training objectives compound this, since overall scene plausibility is weighted far more heavily than a few malformed strokes on a background sign.

That is why lettering deserves to be treated as a defect class rather than a prompting puzzle. Treat garbled lettering the same way you treat warped hands or drifting product geometry, as one more line item in your {{link}}.

Treat garbled lettering the same way you treat warped hands or drifting product geometry, as one more line item in your artifact triage pass.

The clearance trap: invented signage counts as text on screen

The cost is not only aesthetic. In the UK, the ASA and BCAP standards for superimposed text, the small print that qualifies a claim, have been in force since 1 March 2019, and they contain a clause that generative production quietly trips. When you calculate how long a super must be held, all forms of text appearing on screen at any one point in time are counted, including text inside the main ad creative, regardless of where it sits and whether or not it is repeated in the voiceover.

Only a narrow set of text is excluded, essentially the company name, brand name or logo. A generated storefront sign, a product label, a phone screen or a poster in the background is not exempt just because a model invented it. The guidance also warns directly that viewers are less able to read a super when it competes with other on-screen information, and that competing text unrelated to the main message makes comprehension worse.

So a generated plate full of pseudo-lettering does two damaging things at once. It adds text a clearance reviewer may count against your hold calculation, and it degrades the legibility of the one piece of text you are legally obliged to make readable. The frame looks busier, the super gets harder to place, and none of the invented lettering communicates anything to anybody.

A cinematic street scene where all shop signs and product labels are left deliberately blank

What the hold-duration maths costs a generated ad

The arithmetic is published, which makes the cost easy to model before you generate a single frame. Supers are held at a rate of five words per second, or 0.2 seconds per word. On top of that you add a recognition period: two seconds for nine words or fewer, three seconds for ten words or more. Where the qualifying information is particularly significant, practitioners are expected to allow at least a further two seconds.

Take a modest twenty-four-word qualification. That is 4.8 seconds of reading time plus a three-second recognition period, so roughly eight seconds on screen, and closer to ten if the information counts as significant. On a twenty-second spot you have just committed half the runtime to a block of legal text that must stay legible over whatever the model produced behind it. Longer blocks are handled more strictly again: text running to three or more full lines is calculated at a slower words-per-second rate and carries the higher recognition period.

Size and placement are equally prescriptive. Clearcast's guidance for 16:9 HD sets a preferred minimum text height of thirty television lines, dropping to twenty-six when the text sits on an opaque single-coloured block with a margin, and asks for supers to be positioned at the bottom of frame and centred rather than tucked into a corner. Text that qualifies an offer must also appear at the same moment the claim is made, so you cannot solve the problem by pushing everything onto the end card.

An abstract timeline bar divided into coloured blocks representing how long a legal super occupies a short spot

The working rule: generate the plate, set the type in post

The reliable answer is the one conventional production has always used. The camera captures the scene, and a designer sets the type afterwards. Write prompts that keep words out of frame entirely, with no signage, no packaging copy, no screens displaying text and no book covers, because every one of those is an open invitation for the model to invent letterforms.

Then build the type as a real layer, using licensed font files, at full delivery resolution. The type layer belongs in the same {{link}} where you cut, grade and stitch the generated plates.

The payoff compounds quickly. Because the words are live text rather than baked pixels, a legal revision becomes a text edit instead of a regeneration, a second market becomes a translation instead of a new render, and hold duration becomes a timeline adjustment you can make in seconds. Locking fonts, weights and safe-area positions into a documented {{link}} removes the question from every individual shot.

The type layer belongs in the same post-production assembly where you cut, grade and stitch the generated plates.

Locking fonts, weights and safe-area positions into a documented brand consistency system removes the question from every individual shot.

Stacked translucent planes showing a generated video plate separated from a blank type layer above it

Prompt for negative space, not for words

Prompting still matters, just for the opposite goal. Instead of asking for words, ask for the room to put them: a plain wall behind the subject, a defocused background on one side, clean sky above the horizon line. Reserving that negative space at the prompt stage is far cheaper than carving it out of a finished plate, and it stops the composition fighting the lower third later.

Check framing and movement with short, low-resolution renders before paying for a final pass, because a composition that fails at low resolution fails identically at high resolution and costs considerably more to discover there. When you do commit, watch what the delivered file does to fine strokes. Compression and shallow bit depth show up as softness around thin type and as banding in exactly the gradients that white text tends to sit on.

It also helps to reduce the number of stitched joins in the cut. Edits that exist because two generations had to be joined, rather than because an editor chose them, make it harder to time when a super lands and how long a lower third gets to breathe.

The pre-clearance checks that catch this

None of this survives as tribal knowledge, so turn it into a gate. Add a legibility and hold-duration line to your {{link}} so nothing reaches clearance untested.

Four checks catch almost everything. First, scan every generated plate for invented lettering and remove, replace or defocus it. Second, measure text height against a broadcaster test card rather than eyeballing it on a laptop. Third, run the hold duration through a calculator and confirm the qualifying text sits on screen at the same time as the claim it qualifies. Fourth, check contrast against the actual moving background rather than a single frame grab.

The same discipline travels well outside the UK. The FTC's long-standing position on digital advertising is that a disclosure must be clear and conspicuous, placed close to the claim it qualifies, and displayed long enough for an ordinary consumer to notice, read and understand it, which is impossible to argue when the frame behind it is full of visual noise. Contrast ratios and caption styling also overlap heavily with {{link}}, so it is worth solving both in the same pass.

Add a legibility and hold-duration line to your pre-delivery QC checklist so nothing reaches clearance untested.

Contrast ratios and caption styling also overlap heavily with accessibility requirements, so it is worth solving both in the same pass.

Put the framework into production

These related pages connect the article’s planning advice to a specific commercial scope.

Short-form ad productionTurn hook strategy into platform-ready creative variants.AI UGC productionBuild creator-style openings into a controlled testing system.

References

  1. Guidance on superimposed textAdvertising Standards Authority (ASA) / CAP

    The ASA and BCAP's tougher standards for superimposed text came into force on 1 March 2019, requiring the small print in ads to be legible and clear; the guidance counts all forms of text appearing on screen at any one point in time toward the duration of hold, excluding only items such as a company name, brand name or logo, and holds supers at five words per second plus a recognition period of two seconds (nine words or fewer) or three seconds (ten words or more).

  2. Legal Supers - Our Helpful ToolsClearcast

    Clearcast's implementation of the BCAP superimposed text guidance sets a preferred minimum text height of 30 television lines for 16:9 HDTV (26 lines when the text sits on an opaque single-coloured block with a margin), calculates hold at 0.2 seconds per word plus a 2 or 3 second recognition period, applies 0.25 seconds per word plus 3 seconds for supers of three full lines or more, allows up to 5 seconds of recognition time for significant information, excludes brand names from the hold calculation, and requires text qualifying an offer to appear on screen at the same time as the claim, positioned bottom-centre rather than in the corners.

  3. .com Disclosures: How to Make Effective Disclosures in Digital AdvertisingU.S. Federal Trade Commission

    The FTC requires disclosures in digital advertising to be clear and conspicuous, placed as close as possible to the claim they qualify, and displayed prominently and long enough for ordinary consumers to notice, read and understand them.

Related reading

Fixing AI Video Artifacts in Post: A Production PlaybookEditing AI-Generated Video: Turning Loose Clips Into a Finished CommercialAI Video Brand Consistency: The Control Map for Every Brand ElementThe AI Video QC Checklist: Five Gates Before a Cut ShipsAI Video Accessibility: Design the Access Layer Into the Brief