Week of August 01, 2026
The new economics of AI visual production
Seedance 2.5, MiniMax H3 and Veo 3.1 moved AI video from generating shots to directing sequences.
From a clip to a sequence
Until recently these systems were clip generators. You typed a description, received a few seconds of footage, and started again when the character, the product or the camera move came out wrong. The newest releases behave more like a production environment. ByteDance describes Seedance 2.5 as built for thirty-second storytelling with precise reference control and powerful editing capabilities. That claim is about workflow. Picture quality is a separate question.
What direction looks like in practice: reference images and footage go in on the left, the prompt sits underneath, and the generated sequence comes out on the right. The operator specifies physics, framing and continuity instead of describing a scene and hoping for the best.
What the leading models actually produce
Benchmarks rank these systems within a few points of each other. What separates them in practice is what each one is built to be good at, which is easier to show than to describe.
Seedance 2.5
One thirty-second generation holding the same performer, costume and firelight across every cut. Length and continuity in the same shot is the thing that was missing.
ByteDance Seed, Seedance 2.5 launch material
Veo 3.1
Photographic realism carrying weather, skin and cloth at once. The most expensive model on the board at US$24 a generated minute, and it shows in the surfaces.
Google DeepMind, Veo showreel
Gen-4
The same character recognisable across a dozen separate shots and lighting conditions. Recurring people and places are what Runway builds around.
Runway, Gen-4 launch material
These are the vendors' own showreels, chosen because each publishes real sample generations. They show what a system does on material its makers chose. Your own products and characters are a harder test.
Quality converged. Price did not.
On blind preference the leading text-to-video models are now separated by a rounding error. On price they are separated by four times.
US$ per generated minute
4x
The vertical axis is a preference score built from head-to-head votes, and the entire field sits inside about 150 points of it. That is close enough that no provider holds an uncontested lead, so the question stops being which model is best and becomes which one suits the job.
What this does not measure
Blind preference between clips, so nothing about brand accuracy or whether a model holds up in production. Several Chinese entries are variants from one provider, not nine companies. Six of a longer board are shown; Veo 3.1 sits eleventh on it, and Seedance 2.5 is not yet ranked, so 2.0 is plotted. Artificial Analysis, 31 July 2026.
The flywheel matters more than any single model
China had 1.099 billion online audiovisual users at the end of 2025, watching an average of 201 minutes a day, and produced more than two billion AI-generated clips over the year, fourteen times the year before. It is a large audience and an unusually fast one. Feedback arrives in days, so lessons compound.
Micro-dramas, the one to two minute vertical serials made for phones, are the clearest worked example. A five to ten person team can finish an AI-generated series in two to four weeks for around RMB 100,000, against RMB 300,000 to 500,000 and two to three months the traditional way. The AI comic and animated category alone was estimated at RMB 16.8 billion in 2025. None of this makes the drama better. What it buys is the ability to try more things and stop the failures sooner.
Why it drifts, and what now holds it steady
Products, packaging, faces and visual style have to survive from shot to shot and market to market. Consistency here is an operational requirement, and it is what separates a demonstration from a production line.
A demo has to work once. A production system has to work every time.
Why the product or the face changes
What holds it steady
No single provider wins everything, and the strongest setups do not try to make one. On current vendor positioning Seedance 2.5 leads on longer reference-controlled generation, Veo 3.1 on character and scene controls, Runway Gen-4 on recurring visual worlds, and Adobe Firefly on brand-specific consistency. Those are the vendors' own claims, so use them to shortlist what you test, then go and test it.
The number to manage is cost per approved asset
Price per generated second is the figure vendors compete on and very nearly the least useful one. A cheap model becomes expensive when most of what it makes is thrown away.
Everything you spend, divided by the assets you actually approved.
Thirty product videos
Generate 150 alternatives, approve 30. Short, repetitive, and supported by product assets you already own.
HK$3,200
per approved asset
HK$10,000
the old external rate
This is where the saving is real and large. The format is repeatable and the approval risk is low.
One campaign film
AI takes the moodboards, storyboards, location exploration, some B-roll and the market adaptations. Live production, performers, rights and finishing stay.
HK$1.1m
hybrid budget
HK$1.5m
traditional budget
The saving is real but modest. It comes from faster alignment, fewer reshoots and more assets out of a single shoot.
Both scenarios are worked illustrations from this week's brief rather than measured results from a specific company. The point is the shape of the numbers: generation is the smallest line, and the saving depends entirely on which kind of work you are doing.
Cheaper to make, harder to matter
Ipsos and Syracuse University showed ten AI-produced and ten human-produced advertisements to 3,000 US consumers, and Kantar tested AI advertising at scale. The results split cleanly into what AI has already solved and what it has not.
Personalise relevance, not appearance
Cheap production makes it easy to produce a thousand versions. Changing a name or a skyline creates novelty. Changing why the message matters to the person receiving it creates relevance, and only one of those is worth the money.
Generate
Substantially different content per individual or per impression. Held back by cost, privacy, brand safety and who is accountable for it.
Still emerging
Invite
The customer volunteers a photo, a name, an occasion or a preference, and gets something built around it inside boundaries you set.
Highest engagement per version
Segment
Different products, benefits and scenarios for defined audience groups. The creative idea holds; the argument changes.
Where most of the lift is
Adapt
Language, format, market, season and platform, with one core idea underneath. Unglamorous and the fastest thing to get right.
Deployable now, at volume
Mars DINE asked customers to upload a photograph of their cat and describe its behaviour, then generated a video interpreting what the cat might be thinking. Amazon reports 9,500 personalised videos, a 99% completion rate and units up 73% over the six-week campaign against the same period a year earlier. Those are advertiser and platform figures rather than an independent controlled study, but the principle holds: the strongest personalisation reflects what the customer cares about, not what the company knows about them.
SCALE: deciding what to industrialise
Five questions to put to any visual workflow, and the four answers they lead to. Most organisations will find work in all four categories, which is the point.
Which leads to one of four decisions
This is also where the agency relationship gets redrawn. Repetitive production moves closer to the organisation while agencies concentrate on the central idea, the cultural read and the work carrying the most reputational weight. What moves in-house is volume. The judgment stays where it was.
What to do about it
When everyone can produce more content, advantage belongs to the organisation that can preserve identity, direct creativity, earn attention and learn fastest. Four moves that get you there.