Why the last mile is where AI video work fails

Generation stopped being the hard part some time ago. An AI video QC checklist matters now because the bottleneck has moved downstream: a team can render forty usable clips in an afternoon and then discover, three hours before a client call, that the hero shot has a hand which reorganises itself mid-gesture. The failure is rarely the model. It is the absence of a named gate between the render queue and the delivery folder.

The standard the work is judged against moved too. Cannes Lions introduced an AI Craft subcategory for 2026 and set the bar at work whose core concept, execution or impact would not have been possible through previous methods alone, spanning the Design, Digital Craft, Film Craft, Industry Craft and Creative Data Lions. That is a craft threshold, not a novelty allowance.

The decision about choosing between AI and traditional production happens at the top of a project, but the cost of getting it wrong lands here, in the final review. A checklist does not make the work good. It makes failure modes legible, so that a vague note like this feels off becomes shot four breaks object permanence at seven seconds, which is something a producer can actually action.

Gate one: continuity and physics

Watch the cut twice. The first pass runs at full speed with no pausing, because that is how the audience will meet it, and it catches whatever reads as wrong before the analytical brain engages. The second pass is frame-adjacent: pause on every transition, compare the last frame before the cut with the first frame after, and confirm that objects, hands and light sources survive the boundary.

The failures worth naming specifically are morphing edges on hair, fingers and reflective surfaces; objects that disappear for a frame and return; lighting direction that drifts inside a single shot because frames were generated rather than lit; and motion that is mathematically smooth but carries no weight, where people walk without pushing off the ground and objects settle instead of falling.

Log each one with a timecode and a verdict rather than an adjective. Motion break at seven seconds, product edge changes shape during the zoom, acceptable for an internal cut but not for paid media tells the next person exactly what to do. Looks weird does not. Hold the opening seconds to a harder standard than the rest, because an artefact in the first shot spends the viewer's attention before the idea has arrived.

Diagram of sequential film frames on a timeline with misaligned edges highlighted at a cut boundary

Gate two: identity, wardrobe and brand consistency

Continuity inside a shot is a model problem. Continuity across shots is a production problem, and it is what separates a sequence from a pile of clips. Build a reference sheet before generation and audit against it afterwards: face, hair length, wardrobe details, props, the exact product variant, and the colour temperature of the environment.

Product work needs a tighter tolerance than environment work. A background extra whose jacket shifts shade between cuts is survivable. A label that reads correctly in shot two and garbles in shot five is a recall-grade error. Set the rule before review starts, so hero elements are zero tolerance, secondary elements get a judgement call, and background elements are noted and moved past.

Style drift between generated and live-action material is the other half of this gate. When a sequence mixes the two, the generated shots usually arrive glossier and higher in contrast than the plate footage sitting next to them. The fix is a grade pass that pulls both into one world, not another generation cycle.

Gate three: audio, the fastest tell

Audiences forgive a great deal visually and almost nothing sonically. Start with the ambient bed, because a real room is never silent and a scene carrying voiceover and music with no room tone reads as synthetic even when the picture is clean. Then the voice: listen for breath, pacing and emphasis, and flag any delivery that sits flat across a whole sentence.

Music fit is the third check and the most often skipped. Tracks are frequently selected by keyword rather than by feel, so a piece can be technically appropriate and emotionally wrong for the edit. If the beat fights the cut points, the whole thing feels rough even when every frame is clean.

Captions belong in the audio pass rather than the visual one, because caption errors are transcription errors. Keep a canonical terminology list per client covering product names, spellings and shorthand, then check against it. Automatic captions mishandle jargon and branded vocabulary reliably enough that reviewing them should be the default rather than the exception.

Gate four: provenance, disclosure and the paper trail

This is the gate most teams skip and the one that has hardened fastest. YouTube requires creators to disclose generative AI content that makes a real person appear to say or do something they did not do, that alters footage of a real event or place, or that generates a realistic scene which never occurred. The platform also applies labels automatically to material carrying C2PA metadata, which means the disclosure decision can be made on your behalf.

C2PA, the Coalition for Content Provenance and Authenticity, publishes Content Credentials as an open technical standard that lets publishers, creators and consumers establish the origin and edit history of digital content. TikTok became the first video sharing platform to implement it, reading Content Credentials to automatically label AI-generated material uploaded from other tools. Provenance metadata now travels with an asset whether or not anyone planned for it.

Disclosure and rights are separate questions, and the paper trail has to answer both, which means recording the licence tier the footage was generated under alongside whatever label you applied. Keep a per-asset source trail covering the tool, the model version, the reference or prompt used, the licence, the date and the approver. Store it with the approved export rather than in somebody's inbox.

A media file card connected to a chain of small metadata tags representing its production record

Gate five: delivery specs and the sign-off record

Technical delivery is the least interesting gate and the one that generates the most rework. Confirm aspect ratio and safe areas for each destination, caption placement inside those safe areas, the cover frame, loudness normalisation, and the export settings the client's platform actually accepts. A master that has to be re-exported after approval quietly invalidates that approval.

Reviewers should have the brief the work was commissioned against open in a second window, because a technically clean cut that answers the wrong question is still a reshoot. Check the claim, the offer, the legal line and the call to action against approved copy rather than against memory.

Then record the decision. Keep the reviewed export, the completed checklist, the named approver and the timestamp together in one location. This is not bureaucracy. It is what lets a team answer why a particular version shipped six weeks later, and what stops an unapproved cut going out because two files had similar names.

Running the AI video QC checklist without slowing the team down

A checklist nobody uses is worse than none, so tier it. Low-risk internal and organic content gets a two-minute pass: no visible hand or face deformation, audio intelligible, captions correct, disclosure set. Paid media and client-facing work gets all five gates with a named owner on each. Anything featuring a real person, a regulated claim or a hero product gets a second reviewer.

Set escalation thresholds in advance so nobody debates the obvious under deadline. A caption typo in an organic clip is a fix in place. A misstated statistic, an unlicensed track, a warped logo or a missing disclosure is a hard stop regardless of the schedule. Writing these down removes the argument from the moment when there is no time to have it.

Treat the document as living. Every time something slips through, the failure pattern goes into the template so it gets caught by process rather than by whoever happens to be paying attention that week. A cut that clears the gate today still ages, so feed the outcome back into the refresh cadence you set for the campaign.

Schematic of a review funnel splitting a stream of video clips into a fast lane and a five-stage review lane

References

  1. Disclosing use of GenAI contentYouTube Help

    YouTube requires creators to disclose generative AI content that makes a real person appear to say or do something they did not do, alters footage of a real event or place, or generates a realistic scene that did not occur, and may automatically apply an AI label to content containing C2PA metadata.

  2. C2PA: Advancing digital content transparency and authenticityCoalition for Content Provenance and Authenticity

    C2PA provides an open technical standard called Content Credentials that lets publishers, creators and consumers establish the origin and edits of digital content.

  3. Partnering with our industry to advance AI transparency and literacyTikTok Newsroom

    TikTok became the first video sharing platform to implement C2PA Content Credentials, enabling it to read that metadata and automatically label AI-generated content uploaded from other platforms.

  4. Changes for Cannes Lions 2026Cannes Lions

    Cannes Lions introduced an AI Craft subcategory for 2026 requiring entries to demonstrate that the core concept, execution or impact would not have been possible through previous methods alone, appearing across the Design, Digital Craft, Film Craft, Industry Craft and Creative Data Lions.

Related reading

AI TVC vs. traditional production: where each winsAI Video Commercial Rights: How to Keep Client Work SafeHow to Write a Creative Brief for AI Video That Actually DeliversCreative Refresh Cadence: How to Catch Ad Fatigue Before It Drains Budget