Skip to main content

The Video Agent Workflow War: From Generation to Distribution

For the past six months, multi-agent collaboration has become the default for AI video platforms. Services like LibTV and Flova wrapped model APIs into complete script-to-final-cut workflows and quickly scaled their user bases. Now, in late August and early September 2026, the upstream model vendors are joining the game themselves — Alibaba, MiniMax and ByteDance all launched video agent tools within a single month, moving the center of competition from raw model capability to production tools and experience.

Why the battle shifted to workflows

When model capability is no longer scarce, whoever controls the full pipeline from production to distribution will take the lead in the next wave of content industrialization. That is the core judgment driving today's AI video market.

The immediate reason vendors are building tools themselves is data outflow. ByteDance's Seedance holds more than 80% of China's AI video market and reaches 95% penetration in short-drama production. But as a model provider, ByteDance only captures API call volume — how creators break down scripts, adjust storyboards and reuse assets, the truly valuable vertical training material, all stays inside third-party platforms. That material is the core fuel for multimodal model iteration. Vendors are building tools to pull production workflows and data back into their own ecosystems.

Alibaba: five agents in one pipeline

On August 31, Qwen's Agent Teams opened to all users — a full pipeline of screenwriter, director, art director, cinematographer and editor agents. It packs the full Wan 3.0 Prime and Qwen-Image 3.0 Pro into a multi-agent system designed around film-crew roles: a user brings an idea, the screenwriter writes the script, the director storyboards, the art director sets the style, the cinematographer produces frames, and the editor assembles the final cut — all closed inside one platform.

The timing is telling. Wan 3.0 launched on August 24 with 30-second single-shot generation, native 4K resolution and multi-format document input, earning praise in public beta for being "stable, realistic and textured." Although Wan 3.0 was immediately integrated into Flova and LibTV, production workflows and data stayed on those third parties — Alibaba could neither build its own creator ecosystem nor package its model capability into a distinctive production experience. Agent Teams is designed to close exactly that gap.

MiniMax: a workbench built on open source

MiniMax is taking the same path, but its open-source foundation makes the application layer more urgent. The H3 video model, open-sourced in July, costs less than a third of mainstream rivals and performs well in TVC advertising and e-commerce video. But open source cuts both ways: it penetrates the market fast while letting anyone fine-tune and deploy it. Without an application layer, creators struggle to tell H3 apart from competitors, and the market slides into a price war.

MiniMax Design, launched on August 20, is the direct response. It starts four agents — copy, image, video and audio — in parallel, packaging H3's editing, layout, transitions and scoring into out-of-the-box workflows. Notably, its presets are not simple templates; they are style-calibrated to H3's attention mechanism, making the "native model understands content better" advantage real in production. On the same day Agent Teams launched, MiniMax announced H3 Max integration with Design — generating a full 5-second 768P audio-video clip in under 3 seconds, faster than playback speed, opening the door to live-streaming scenarios.

ByteDance: defending dominance with a closed loop

For ByteDance, the market leader, competition has spread from model parameters to production tools and creator ecosystems. Its comic-drama creation tool, in beta since August 18, covers the full pipeline from script upload to video assembly, supports one-click conversion of scripts up to 200,000 characters, allows 30 people online in one team space, and binds directly to Douyin's short-drama creator center — connecting production and distribution data. A day later, Seedance Studio entered small-scale invite testing with project-based management, targeting professionals who want to "make a film" rather than "generate a clip."

The strategic fulcrum is linking ByteDance's three cards — model, tool and distribution — into a closed loop: upstream draws on the massive IP library of Tomato Novel, midstream handles full AI collaboration from script breakdown to video assembly, and downstream syncs to the Douyin short-drama creator center for publishing, review and revenue sharing.

ByteDance also plays a dynamic pricing game. While Seedance 2.5 pushes the high-end quality ceiling, the 2.0 mini and fast tiers dropped in price: 2.0 mini at 720P down to a limited-time 0.2 yuan/second for low-cost experimentation, and 2.0 fast at 720P to 0.6 yuan/second, directly targeting MiniMax H3's pricing band. Premium products build the brand, mid-tier products fight for share, and entry products expand the user base.

Aggregators: squeezed but not obsolete

The pressure falls hardest on third-party aggregators like LibTV and Flova. August's price war reached absurd levels — even accounting for resolution differences, the same Seedance 2.5 tier can vary several-fold across platforms. When every platform runs the same upstream models, products can't differentiate, and competition collapses into access speed and subsidies.

Yet aggregators keep two irreplaceable values. First, they are the vendors' most important distribution channel — Wan 3.0 joined Flova, LibTV and RunningHub on day one, and MiniMax H3 chose LibTV for its debut. Second, multi-model mixing gives them room to survive: mixing Seedance, MiniMax and Kling models is nearly consensus among AI short-drama practitioners, while official vendor agents naturally favor their own models — a real constraint in complex cross-model projects.

Aggregators are also building moats: LibTV invested 10 million yuan to support Skill creators and distill top directors' workflows into SkillHub; OiiOii launched 149 preset style libraries with a global asset-lock system; TapNow uses open-source workflow libraries for distributed teams with strict version control.

What it means for the industry

This is the moment AI video shifts from "generation capability" to "industrial capacity." When model capability stops being scarce, pipeline capability becomes the new dividing line — whoever turns a model into a competitive industrial production tool will hold the content-industry advantage from production to distribution. For practitioners, understanding multi-agent workflows, cross-model mixing and full-pipeline loops is becoming basic literacy. The migration from models to workflows signals that the AI content industry is entering a new phase built on tools, ecosystems and distribution.