Live event video production just automated the one job that had no post-production
Live event video production is the last video workflow where post-production never existed, and in 2026 that changed: reframing and clip generation moved inside the live encoder. AWS Elemental Inference, launched in MediaLive on 24 February 2026, crops a landscape feed to vertical and extracts highlight clips in 6 to 10 seconds, running in parallel with the encoder and requiring no human-in-the-loop prompting. The edit decision now happens before the stream finishes, so it has to be specified before the event begins.
Until this year the workaround was a static centre crop. A vision mixer framed wider than the shot deserved so the 9:16 version would not lose a presenter, and the vertical audience still watched half a body the moment anyone walked. Highlight clips were a post-event task that shipped hours after the moment had stopped mattering. The latency that once separated a live workflow from a rendered one is now counted in seconds, which is the same shift behind {{link}}.
What separates this from a smarter crop tool is where it runs. The service analyses the feed in parallel with the encoder rather than after it, applies several features to the same stream at once, and returns 6 to 10 seconds of latency against minutes for a post-processing pass. AWS describes the pattern as process once, optimize everywhere. The consequence is straightforward: the crop and the clip are produced while the event is still running, so no reviewer stands between the model and the audience.
The latency that once separated a live workflow from a rendered one is now counted in seconds, which is the same shift behind the shift into real-time generation.
What the live encoder now decides without asking
At launch the service does two jobs. Vertical video creation crops a landscape broadcast into 9:16 while tracking subjects and keeping the key action in frame. Clip generation detects and extracts moments from live content for immediate distribution; AWS illustrates it with game-winning plays in soccer and basketball and frames the gain as reducing manual editing from hours to minutes. The two detections run independently, and the documentation states plainly that the service applies its AI capabilities with no human-in-the-loop prompting required.
Two consequences have to be absorbed before the first whistle. The first is that editorial judgment moves upstream. A tracking crop needs to know who the subject is, and a clipping model needs to know what counts as a highlight, which means both have to be expressed as references and rules before the event starts rather than corrected during it. The second is that the foundation models are fully managed and automatically updated. A crop that behaved one way last season is not guaranteed to behave that way next, and nothing in the workflow forces that change to reach the people planning the next broadcast.
Availability is the third constraint, and it is the one that quietly decides whether any of this reaches a given production. The service launched in four AWS Regions, with consumption-based pricing and no upfront commitment, which suits a tournament or a product launch far better than a permanent channel. A regional event in a market those four regions do not cover cannot use the pattern at all, and a production that wants it has to route its contribution feed accordingly.

Why source resolution became a delivery decision again
A 1080p landscape frame is 1920 pixels wide and 1080 tall. A vertical canvas is 1080 by 1920, which is taller than the source itself, so every vertical output cut from it is either enlarged or delivered below 1080 lines. Ingest the same event at 3840 by 2160 and the vertical frame is sampled natively at full width, because 1080 by 1920 fits inside the source with room to spare. A resolution decision that used to concern the broadcast master now determines the quality of every derivative the encoder publishes.
The trade is not free. YouTube's live encoder settings call for 10 to 40 Mbps for 2160p at 60 frames per second using AV1 or H.265, and 35 Mbps with H.264, against 4 to 10 Mbps for 1080p at the same frame rate, with a keyframe every two seconds and CBR encoding. The same page notes that selecting 4K/2160 disables low-latency optimisation. A production that wants natively sampled vertical output is therefore buying bandwidth and giving up latency, which is the same logic that runs through {{link}}, where containers and grades stopped being a finishing step and became part of what gets delivered.
The alternative used to be a second camera and a second crew, or a delayed post-event pass that turned vertical clips around the following morning. Both of those options at least kept the decision with an editor. Once the vertical frame is derived from the same encoder that makes the broadcast, framing stops being a separate job with its own budget line and becomes a parameter of the ingest, and the person who signs off on that parameter is usually the technical director rather than the creative lead.
A production that wants natively sampled vertical output is therefore buying bandwidth and giving up latency, which is the same logic that runs through AI video HDR delivery, where containers and grades stopped being a finishing step and became part of what gets delivered.

Clips that exist before the moment ends
Clip generation does not stay inside the broadcaster. On 10 July 2026 TikTok and WSC Sports announced a partnership that routes AI-cut vertical clips to vetted creators. Magicrop reframes a clip into vertical with the action tracked in frame, ready to publish in minutes, drawing on live action, archive material and behind-the-scenes footage pooled into one searchable library. Creators receive licensed content they could not otherwise clear, while the rights holder keeps control of usage, rights and brand safety. WSC Sports says its platform serves the NBA, ESPN, LaLiga and roughly 650 organisations in total, a figure the company states itself.
The interesting part is not the automation but where control now sits. Rights holders historically resisted full short-form distribution because cutting clips at the volume a live calendar produces meant either a large in-house editing team or accepting fan clips nobody could control. Routing automated cuts through pre-approved creators answers the second problem without solving the first, and it turns approval rather than editing into the scarce resource. Live sport was the first proving ground for {{link}}, and the same distribution machinery now reaches events with a fraction of that audience.
Underneath that sits a metadata problem nobody wants to own. A clip is publishable only if the system already knows the territory it applies to, the window in which it may run, whether a sponsor mark is permitted inside it, and whether the people in frame are cleared for that market. Those rules usually live in contracts and inboxes rather than in a machine-readable form, so an automated clipping layer will either move more slowly than the event or publish something it should not.
That leaves a governance question with nothing to do with model quality. When a crop misses the point of a play, or a clip is cut a beat early, the mistake is already in front of an audience. A human editor who cuts badly is corrected before publication; an automated one is corrected afterwards, in public, on the rights holder's account.
Live sport was the first proving ground for live sports video advertising, and the same distribution machinery now reaches events with a fraction of that audience.

What to specify before the event, not during it
The list is short and unglamorous. Define the crop subject per camera, because a tracking crop inherits whatever the operator treats as the point of the shot. Define what counts as a clip, with minimum and maximum lengths and a margin of seconds before and after the moment, or the clipping model will decide for you. Justify the ingest resolution against the vertical derivative rather than the master. Name who approves a clip and the window in which that approval has to happen. Record the model version behind any output you keep, since managed models update without asking. And keep a provenance record for every derivative, because a crop is a new asset even when the pixels came from your own feed.
The audience arithmetic makes the deadline real. Attest's 2026 US media consumption research puts social video at three hours and 54 minutes a day against three hours and 20 minutes for live television and subscription streaming combined, and finds that only 12% of viewers give television their full attention; about two thirds regularly watch YouTube videos longer than 15 minutes. Live video production is serving those same hours from two formats at once. Getting the specification right is the discipline that defines {{link}}, where delivery stopped being a final handover and became the test of whether work counts as finished at all.
None of this requires a new department. It requires the production team to treat the encoder as an editorial position with a rulebook, rather than as plumbing that happens to output a vertical file. The events that get this right will look identical to the ones that do not, until the first clip published without approval appears on a platform nobody was watching.
Getting the specification right is the discipline that defines the AI video delivery era, where delivery stopped being a final handover and became the test of whether work counts as finished at all.
Put the framework into production
These related pages connect the article’s planning advice to a specific commercial scope.
References
- Transform live video for mobile audiences with AWS Elemental InferenceAmazon Web Services
AWS Elemental Inference launched on 24 February 2026 as a fully managed AI service that adapts live and on-demand video into vertical formats in real time and generates clips from live content. It runs AI features in parallel with live video encoding at 6-10 seconds of latency against minutes for post-processing, detects vertical cropping and clip generation independently, requires no human-in-the-loop prompting, uses fully managed foundation models that are automatically updated, launched in four AWS Regions, and is priced on consumption.
- TikTok and WSC Sports partner to give rights holders a new way to reach fans through content creatorsWSC Sports
Announced on 10 July 2026, the TikTok and WSC Sports partnership routes AI-cut vertical clips to vetted TikTok creators while rights holders retain control of usage, rights and brand safety. WSC Sports' Magicrop turns a clip into vertical video with the action tracked in frame and ready to publish within minutes, combining live action, archive content and behind-the-scenes footage in one searchable platform, and the company states that it serves the NBA, ESPN, LaLiga and around 650 sports organisations.
- Choose live encoder settings, bitrates, and resolutionsYouTube Help
YouTube's live encoder guidance recommends 10-40 Mbps for 4K/2160p at 60fps using AV1 or H.265 and 35 Mbps with H.264, against 4-10 Mbps for 1080p at 60fps using AV1 or H.265, with a keyframe interval of two seconds (no more than four), CBR bitrate encoding and a square pixel aspect ratio; it also states that selecting 4K/2160 does not allow low-latency optimisation.
- 2026 US media consumption reportAttest
Attest's 2026 US media consumption research reports social video consumption at three hours 54 minutes a day versus three hours 20 minutes for live television and subscription streaming combined, finds that 12% of viewers give television their full attention, and reports that around two thirds of people regularly watch YouTube videos longer than 15 minutes.
