Where Video Attention Went in 2026

Video attention did not shrink in 2026 - it changed hands. Social video now takes 3 hours 54 minutes of the average US adult's day, more than live TV and subscription streaming combined at 3 hours 20 minutes. But only 12% of people give television their undivided attention, and forcing the issue does not work: when an ad cannot be skipped, just 44% say they actually watch it.

Attest's 2026 US media consumption report rests on five years of US media-habits tracking plus a fresh, nationally representative survey of 1,000 US adults. Read it as a reallocation rather than a decline. Creator platforms are not stealing spare minutes from television any more; on a daily time-spent basis they have become television, and every media plan still built on the old hierarchy is priced against the wrong benchmark.

Two details matter more than the headline number. Roughly two-thirds of people say they regularly watch YouTube videos longer than 15 minutes, and TikTok - the platform built on 15-second clips - now holds viewers for content roughly 60 times that length. The format credited with destroying attention is where the longest voluntary viewing sessions now happen. Attention did not get smaller. It got choosier about where it goes.

The 12% figure is the one to put in front of a client. It means that for roughly seven in eight television viewers, the screen is one input among several: a phone, a shopping tab, a message thread. The advertiser is not competing with the programme, or even with the other advertisers. The advertiser is competing with whatever is in the viewer's hand.

Bar comparison diagram showing social video time exceeding television and streaming time combined in an average day

Forced Exposure Is Not Earned Attention

The skip control is the most honest measurement instrument in the business. In Attest's data, 86% of Gen Z say they always or usually skip ads, and Boomers do it nearly half the time. Buying a non-skippable slot does not close the gap. When people cannot skip, only 44% say they watch the ad as intended, 34% redirect their attention elsewhere and 17% simply mute the sound. A forced impression is worth somewhere between a third and a half of a voluntary one, and the discount is steepest with the audiences most briefs say they want.

Generative video adds a second problem: it supplies a reason to opt out. Thirty-five percent of Gen Z respondents name AI-generated ads as a reason to disengage. Ninety-two percent of people use at least one ad-supported streaming platform, so the inventory exists in abundance - what is scarce is tolerance inside it. The buy-side version of this is documented in our report on {{link}}.

The arithmetic is unflattering once you apply it. A campaign that reports one million delivered impressions on non-skippable inventory has, on this data, roughly 440,000 actual views, about 300,000 of which involve the viewer looking elsewhere. None of that shows up in a delivery report. It shows up later, as a creative problem, because the creative is the only part of the chain anyone can still change.

The buy-side version of this is documented in our report on generative video in performance media.

Funnel diagram showing a delivered impression narrowing through skipped, redirected and muted viewing to a smaller watched remainder

Cheap Supply Made Attention Scarcer

Generative tools drove the marginal cost of a finished cut close to zero, and the market responded the way markets do - by shipping more of them. Metricool's 2026 TikTok study, covering 2,314,756 posts from more than 92,000 accounts, found videos published up 72.10% year over year while video views per post fell 31.30%, reach fell 28.73% and interactions fell 31.17%. Supply rose; the attention available to any single asset fell. That is the mechanism behind most 'our creative suddenly stopped working' conversations this year.

It also compresses the window. The same study found posts collect 96% of their total reach and nearly 98% of their interactions inside the first 10 days, so an asset has about a week and a half to earn whatever it is going to earn. The duration question compounds it, because our analysis of {{link}} shows the formats that win on engagement are rarely the ones that win on views.

This is the part AI video did not fix. It solved the cost of production and left the cost of attention untouched, so the ratio between what teams can make and what audiences will sit through has widened by roughly an order of magnitude in two years. Every downstream symptom - falling per-post views, shorter asset lifespans, more time spent justifying creative - is a symptom of that ratio.

The duration question compounds it, because our analysis of the short-form video duration paradox shows the formats that win on engagement are rarely the ones that win on views.

What Earned Attention Looks Like in a Cut

Treat the first two seconds as a decision point, not an introduction. Logos, title cards and slow establishing moves are bought at full rate and returned at a fraction of it. Put the claim, the product and the reason to stay in frame before the skip control becomes available, and let brand identification ride along with them rather than precede them. Everything that happens before the viewer has a choice is overhead.

Then assume the sound is off. Seventeen percent of people mute an ad they cannot skip, and far more were never going to turn the volume up in the first place. The proposition has to survive as burned-in text: legible supers, on-screen product naming, and pacing that still reads in silence. Generated lettering is unreliable enough that on-screen text belongs in post rather than in the prompt. Separating a hook failure from a body failure is the first diagnostic step, and {{link}} turns that into a repeatable sequence.

The practical version is a mute-and-first-frame pass. Watch the cut once with the sound off and the opening two seconds frozen. If the proposition does not survive that, no amount of media weight will recover it. Most failures surface in that pass, and the remedy is usually a re-cut rather than a re-shoot: the footage is fine, the sequencing is not.

Separating a hook failure from a body failure is the first diagnostic step, and AI video retention diagnostics turns that into a repeatable sequence.

Storyboard row of four panels showing the opening seconds of a cut carrying the claim with a crossed-out speaker icon indicating silent viewing

Long-Form Is Where Opt-In Attention Lives

The counterweight is that voluntary attention has not collapsed; it has concentrated. Wistia's 2026 State of Video data shows on-demand webinars still pulling views six months after the live event, with replays longer than 30 minutes drawing twice the engagement of shorter ones, and videos over 60 minutes holding a 52% play rate. Homepage video, by comparison, sees a 24% play rate. Set that against a 10-day reach window on social and the strategic picture inverts: the long cut is the durable asset, the short cut is the consumable one.

That only pays if the long cut is navigable. What turns chapters and key moments into entry points a viewer can jump to directly is {{link}}, not another edit pass. When those moments are declared in markup they also take precedence over a platform's automatic detection. Where the asset should then live is a separate decision, and it is the one {{link}} maps out.

The organisational consequence is that in-house teams now carry this work. Wistia attributes the jump from 36% to 54% of companies running in-house video teams squarely to AI, with 62% already using or planning to use AI in video workflows. Cheap generation did not remove the need for an editor. It moved the editor in-house and handed them a much larger volume of material to make watchable.

What turns chapters and key moments into entry points a viewer can jump to directly is video structured data, not another edit pass.

Where the asset should then live is a separate decision, and it is the one the B2B video distribution shift maps out.

A Video Attention Checklist for the Next Brief

Four questions to answer before a cut enters production. Can the proposition be understood with the sound off? Does the first two seconds carry the claim instead of the logo? Is there a long-form version capable of earning attention for months rather than days? And are the chapters declared in the file and its markup, not only in the edit?

None of this needs new tooling. It needs conceding that the impression you bought and the attention you received are different numbers, and that in 2026 the second one is the scarcer commodity. AI video removed the cost constraint on making things. It did nothing to raise the supply of attention available to watch them, and it is too late to pretend the two curves will converge on their own. Briefs, cuts and delivery specs all have to be rewritten around the smaller number, because that is the only one the audience actually controls.

Put the framework into production

These related pages connect the article’s planning advice to a specific commercial scope.

Short-form ad productionTurn hook strategy into platform-ready creative variants.AI UGC productionBuild creator-style openings into a controlled testing system.

References

  1. 2026 US media consumption report: The Attention EconomyAttest

    Social video claims 3h 54m of the average person's day versus 3h 20m for live TV and subscription streaming combined; only 12% of people give TV their full undivided attention; 86% of Gen Z always or usually skip ads; when an ad cannot be skipped only 44% watch it as intended, 34% redirect their attention and 17% mute; 35% of Gen Z cite AI-generated ads as a reason to disengage; 92% use at least one ad-supported streaming platform. Based on five years of US media habits tracking and a nationally representative survey of 1,000 US adults.

  2. Metricool 2026 TikTok StudyMetricool

    Across 2,314,756 posts from more than 92,000 accounts, videos published rose 72.10% year over year while video views per post fell 31.30%, reach fell 28.73% and interactions fell 31.17%; posts collect 96% of total reach and nearly 98% of total interactions within the first 10 days.

  3. The Key Takeaways From Wistia's 2026 State of Video WebinarWistia

    On-demand webinars are still pulling views six months after the live event, replays longer than 30 minutes see twice the engagement of shorter ones, and videos over 60 minutes hold a 52% play rate against 24% for homepage video; in-house video teams grew from 36% to 54% in two years, a change Wistia attributes to AI.

  4. Video structured data: key moments and Clip markupGoogle Search Central

    Key moments declared in structured data through Clip and hasPart take precedence over Google's automatically detected key moments, letting publishers define the entry points into a long video.

Related reading

Generative Video Is Becoming Performance Media: How Brands Turn AI Video Into Measurable ROASThe 2026 Short-Form Video Duration ParadoxReading the AI Video Retention Curve: A Diagnostic Map From Hook to EndingAI Video SEO: A Structured-Data Playbook for Commercial TeamsB2B Video Distribution in 2026: Redistributing AI-Generated Video Beyond YouTube's Collapsing Organic Reach