Preface Executive Briefing

Week of August 01, 2026

The new economics of AI visual production

Seedance 2.5, MiniMax H3 and Veo 3.1 moved AI video from generating shots to directing sequences.

From a clip to a sequence

Until recently these systems were clip generators. You typed a description, received a few seconds of footage, and started again when the character, the product or the camera move came out wrong. The newest releases behave more like a production environment. ByteDance describes Seedance 2.5 as built for thirty-second storytelling with precise reference control and powerful editing capabilities. That claim is about workflow. Picture quality is a separate question.

Prompt and hope

A few seconds at a time, described in words. Anything wrong meant regenerating the whole thing and losing what already worked.

Direct and revise

Thirty seconds in one generation, extended twice, steered with reference footage, and edited in place. Duration, direction, continuity, editability and cost all moved at once.

What direction looks like in practice: reference images and footage go in on the left, the prompt sits underneath, and the generated sequence comes out on the right. The operator specifies physics, framing and continuity instead of describing a scene and hoping for the best.

Excerpt from ByteDance Seed’s Seedance 2.5 launch video.

What the leading models actually produce

Benchmarks rank these systems within a few points of each other. What separates them in practice is what each one is built to be good at, which is easier to show than to describe.

Seedance 2.5

One thirty-second generation holding the same performer, costume and firelight across every cut. Length and continuity in the same shot is the thing that was missing.

ByteDance Seed, Seedance 2.5 launch material

Veo 3.1

Photographic realism carrying weather, skin and cloth at once. The most expensive model on the board at US$24 a generated minute, and it shows in the surfaces.

Google DeepMind, Veo showreel

Gen-4

The same character recognisable across a dozen separate shots and lighting conditions. Recurring people and places are what Runway builds around.

Runway, Gen-4 launch material

These are the vendors' own showreels, chosen because each publishes real sample generations. They show what a system does on material its makers chose. Your own products and characters are a harder test.

Quality converged. Price did not.

On blind preference the leading text-to-video models are now separated by a rounding error. On price they are separated by four times.

Blind-preference rating
1100
1150
1200
1250
Gemini Omni Flash
MiniMax H3
Seedance 2.0
Wan 2.7 (Alibaba)
HappyHorse 1.1 (Alibaba)
Veo 3.1
$6$12$18$24

US$ per generated minute

China-basedUS-based

4x

The vertical axis is a preference score built from head-to-head votes, and the entire field sits inside about 150 points of it. That is close enough that no provider holds an uncontested lead, so the question stops being which model is best and becomes which one suits the job.

What this does not measure

Blind preference between clips, so nothing about brand accuracy or whether a model holds up in production. Several Chinese entries are variants from one provider, not nine companies. Six of a longer board are shown; Veo 3.1 sits eleventh on it, and Seedance 2.5 is not yet ranked, so 2.0 is plotted. Artificial Analysis, 31 July 2026.

The flywheel matters more than any single model

China had 1.099 billion online audiovisual users at the end of 2025, watching an average of 201 minutes a day, and produced more than two billion AI-generated clips over the year, fourteen times the year before. It is a large audience and an unusually fast one. Feedback arrives in days, so lessons compound.

01

Production gets cheaper

Generating a finished minute costs single-digit dollars, with no shoot to organise.

02

More gets tried

Teams can afford formats, niches and endings they would never have funded before.

03

Audiences answer faster

Completion, drop-off, comments and reposts come back within days instead of quarters.

04

The signal sharpens

Failures are killed early and the patterns that hold attention are found sooner.

05

The tools improve

Model builders sit next to that demand, so the next release is shaped by it.

Micro-dramas, the one to two minute vertical serials made for phones, are the clearest worked example. A five to ten person team can finish an AI-generated series in two to four weeks for around RMB 100,000, against RMB 300,000 to 500,000 and two to three months the traditional way. The AI comic and animated category alone was estimated at RMB 16.8 billion in 2025. None of this makes the drama better. What it buys is the ability to try more things and stop the failures sooner.

Why it drifts, and what now holds it steady

Products, packaging, faces and visual style have to survive from shot to shot and market to market. Consistency here is an operational requirement, and it is what separates a demonstration from a production line.

A demo has to work once. A production system has to work every time.

Why the product or the face changes

It has never seen the whole thing

One reference photograph does not show every angle of a product or a face. The moment the subject turns, the model has to invent what it was never given, and what it invents looks convincing while being wrong.

You asked for too much at once

Hold this character, but change the pose, the lighting, the clothing, the room and the camera move. The more that changes around a subject, the less of the subject survives.

Small errors compound

Video has to hold a subject steady across hundreds of frames. A logo shifts, a jacket changes shade, a face slowly becomes a different face. With several people in shot, features start attaching to the wrong person.

What holds it steady

  • More referencesSeveral approved images, angles and clips instead of one. Cuts down what the model has to invent.
  • Tracking across framesThe system follows the subject and its surroundings through the shot, so continuity stops being luck.
  • Identity trainingThe model is rewarded for keeping the same person or product. Protects the assets you cannot afford to get wrong.
  • Anchor framesApproved stills locked in as checkpoints before any motion is generated. Turns a gamble into something predictable.
  • Custom models and targeted editsTrain the system on your own products and characters, then fix one element instead of regenerating everything.
Approved brand keyframe
Specialist generation
Targeted editing
Traditional finishing

No single provider wins everything, and the strongest setups do not try to make one. On current vendor positioning Seedance 2.5 leads on longer reference-controlled generation, Veo 3.1 on character and scene controls, Runway Gen-4 on recurring visual worlds, and Adobe Firefly on brand-specific consistency. Those are the vendors' own claims, so use them to shortlist what you test, then go and test it.

The number to manage is cost per approved asset

Price per generated second is the figure vendors compete on and very nearly the least useful one. A cheap model becomes expensive when most of what it makes is thrown away.

Everything you spend, divided by the assets you actually approved.

Model and platformHK$12,000
Internal directionHK$30,000
Editing and compositingHK$24,000
Brand and legal reviewHK$12,000
Agency finishingHK$18,000
Total for 30 approved videosHK$96,000

Thirty product videos

Generate 150 alternatives, approve 30. Short, repetitive, and supported by product assets you already own.

HK$3,200

per approved asset

HK$10,000

the old external rate

This is where the saving is real and large. The format is repeatable and the approval risk is low.

One campaign film

AI takes the moodboards, storyboards, location exploration, some B-roll and the market adaptations. Live production, performers, rights and finishing stay.

HK$1.1m

hybrid budget

HK$1.5m

traditional budget

The saving is real but modest. It comes from faster alignment, fewer reshoots and more assets out of a single shoot.

Both scenarios are worked illustrations from this week's brief rather than measured results from a specific company. The point is the shape of the numbers: generation is the smallest line, and the saving depends entirely on which kind of work you are doing.

Cheaper to make, harder to matter

Ipsos and Syracuse University showed ten AI-produced and ten human-produced advertisements to 3,000 US consumers, and Kantar tested AI advertising at scale. The results split cleanly into what AI has already solved and what it has not.

Nobody can tell

Most viewers could not confidently say which advertisements were made with AI. On surface quality the gap has closed.

  • Credible production values at a fraction of the cost
  • No penalty for using AI where the work is seamless
  • More than 40% of well-integrated AI ads were both noticed and correctly linked to the brand

Nobody remembers

The human-produced work still won on the things that make an advertisement worth running at all.

  • Stronger on entertainment, uniqueness, talkability and emotional pull
  • 14% stronger on short-term effectiveness, 17% on long-term brand impact
  • Branding fell away when the visuals were left to a generic model

Personalise relevance, not appearance

Cheap production makes it easy to produce a thousand versions. Changing a name or a skyline creates novelty. Changing why the message matters to the person receiving it creates relevance, and only one of those is worth the money.

  1. Emerging

    Generate

    Substantially different content per individual or per impression. Held back by cost, privacy, brand safety and who is accountable for it.

    Still emerging

  2. Bounded

    Invite

    The customer volunteers a photo, a name, an occasion or a preference, and gets something built around it inside boundaries you set.

    Highest engagement per version

  3. Proven

    Segment

    Different products, benefits and scenarios for defined audience groups. The creative idea holds; the argument changes.

    Where most of the lift is

  4. Baseline

    Adapt

    Language, format, market, season and platform, with one core idea underneath. Unglamorous and the fastest thing to get right.

    Deployable now, at volume

Mars DINE asked customers to upload a photograph of their cat and describe its behaviour, then generated a video interpreting what the cat might be thinking. Amazon reports 9,500 personalised videos, a 99% completion rate and units up 73% over the six-week campaign against the same period a year earlier. Those are advertiser and platform figures rather than an independent controlled study, but the principle holds: the strongest personalisation reflects what the customer cares about, not what the company knows about them.

SCALE: deciding what to industrialise

Five questions to put to any visual workflow, and the four answers they lead to. Most organisations will find work in all four categories, which is the point.

S

Standardisable

Is the work recurring and structured enough to be worth a system?

C

Controllable

Can approved products and characters genuinely constrain the output?

A

Acceptance yield

What share of what comes out is usable without heavy repair?

L

Liability

Are protected identities, regulated claims or hard-won trust in play?

E

Economics

Does it improve cost per approved asset and the business result?

Which leads to one of four decisions

Industrialise

Recurring, controllable, low consequence if a version is imperfect. Build the pipeline and run it.

Localisation, platform adaptations, product variants, internal video

Co-produce

AI earns real efficiency but the craft still decides whether it works. Keep human direction on the critical path.

Major campaigns, spokesperson programmes, high-value launches

Experiment

Credible, but not dependable yet. Fund it as a test with a real question attached instead of a rollout.

One-to-one generated video, longer AI-native brand entertainment

Protect

Being wrong here is not recoverable by reshooting. Keep it human and keep it reviewed.

Executive communication, regulated claims, unlicensed likenesses

This is also where the agency relationship gets redrawn. Repetitive production moves closer to the organisation while agencies concentrate on the central idea, the cultural read and the work carrying the most reputational weight. What moves in-house is volume. The judgment stays where it was.

What to do about it

When everyone can produce more content, advantage belongs to the organisation that can preserve identity, direct creativity, earn attention and learn fastest. Four moves that get you there.

01

Measure your current cost per approved asset first

Almost nobody knows it. Without that baseline every saving claimed from AI is a guess, and the first honest number usually reveals that generation was never the expensive part. Track acceptance yield and time to approved cut alongside it.

02

Turn brand guidelines into production infrastructure

Guidelines were written for people. A model needs approved products from several angles, characters, environments, mandatory details, prohibited treatments and rights metadata. The model's general knowledge is available to every competitor; your own assets are not.

03

Redesign approval before you increase volume

Production capacity is now easy to buy and approval capacity is not. A content factory without a matching approval factory produces a queue, and the review bottleneck is where the promised savings quietly disappear.

04

Measure attention, not output

Track distinctiveness, brand attribution, emotional response and incremental sales beside cost and speed. The goal is more valuable attention for every creative and media dollar.

FAQ

Seedance 2.5 is ByteDance's latest video model. Thirty seconds matters because it is long enough to hold a complete idea. Below that you are stitching disconnected clips together and hoping the character, product and lighting survive the joins. At thirty seconds, with two further extensions and the ability to feed in reference footage, the work becomes direction instead of prompting.