Preface Executive Briefing

Week of July 11, 2026

The week the US AI race changed shape

OpenAI targets the work. xAI targets the cost curve. Meta targets a comeback.

The signal

Three releases. Three different ways to compete.

Within two days, OpenAI, xAI and Meta each gave a different answer to the same question: how do you win when every lab has a strong model?

OpenAI8 to 9 Jul
Work

Make AI the place where work gets finished

GPT-5.6 went live alongside ChatGPT Work, the GPT-Live voice model and a unified desktop app. Microsoft 365 Copilot now lists GPT-5.6 as its preferred model.

xAI8 Jul
Economics

Make frontier intelligence cheaper to run

Grok 4.5 reaches near-frontier benchmark scores at $0.31 per measured task, using far fewer tokens than the models it competes with.

Meta9 Jul
Reach

Make switching easy and distribution hard to match

Muse Spark 1.1 arrives with Meta's first paid API, priced at roughly a quarter of rival flagship rates and compatible with the request formats developers already use.

OpenAI's move

From the smartest model to the finished work

GPT-5.6 splits OpenAI's flagship into three durable tiers: Sol, Terra and Luna. You choose how much intelligence, and how much cost, each task deserves.

GPT-5.6 SolFor the hardest professional work59Intelligence Index$1.04 per task
GPT-5.6 TerraFor everyday work55Intelligence Index$0.55 per task
GPT-5.6 LunaFor speed and volume51Intelligence Index$0.21 per task
One portfolio, many places to finish work
SlidesSpreadsheetsDocumentsWebsitesChatGPT WorkVoice via GPT-LiveMicrosoft 365 CopilotGitHub Copilot

Scores and task costs are Artificial Analysis measurements at each model's highest reasoning effort. The strategy sits around the models: the same week brought ChatGPT Work, the GPT-Live voice model, a unified desktop app and preferred placement in Microsoft 365 Copilot. OpenAI is selling the finished work, not the chat window.

xAI's move

The breakthrough is the cost of the answer

Grok 4.5 comes from xAI, which folded into SpaceX and rebranded SpaceXAI this month. Its headline result is not the score. It is what each completed task costs.

Intelligence
45
50
55
60
GPT-5.6 Sol
GLM-5.2
Grok 4.5
Kimi K2.6
GPT-5.6 Luna
$0.25$0.50$0.75$1.00

US$ per completed Intelligence Index task

Grok 4.5 (high reasoning)US frontier modelsChinese challengers

$0.31

per benchmark task

Artificial Analysis measured Grok 4.5 at $0.31 per Intelligence Index task: near-frontier intelligence at the task price of Kimi K2.6, which scores 11 points lower. Price is only half of it. Grok needs about 14,000 output tokens per task, far fewer than its rivals.

Cheap is not the same as reliable

On the same evaluations, Grok 4.5's hallucination rate rose to 54%, from 25% for Grok 4.3, even as accuracy improved. Budget for reliability testing before routing real work to it.

Artificial Analysis Intelligence Index at each model's tested reasoning configuration, retrieved 11 July 2026. Measured benchmark cost, not a production quote.

Meta's move

A comeback built on price and reach

Muse Spark 1.1 is Meta's first paid, proprietary model. That is a deliberate break from the open-weight Llama strategy that built its developer following.

Model

Muse Spark 1.1

A million-token agentic model for coding, computer use and orchestrating parallel agents.

API

Priced to try

$1.25 in and $4.25 out per million tokens, and it accepts OpenAI-style and Anthropic-style requests.

Distribution

Meta's platforms

Free consumer access in the Meta AI app, plus products billions of people already use.

unproven
Adoption

Still to be earned

Independent evidence is thin so far. Where it exists, it runs below Meta's own numbers.

Meta reports a Terminal-Bench score of 80.0. Independent evaluator Vals AI measured 69.29. Artificial Analysis has not yet published results for Muse Spark 1.1.

Treat the comeback as an attempt, not a result. It can still move your market: a cheap, compatible API attached to Meta's distribution forces rivals to respond.

The new model race

Work. Economics. Reach.

Put the three moves on one map and the week's real story appears: the advantage now sits around the model as much as inside it.

WorkOpenAI

Own the finished output and the surfaces where work happens.

EconomicsxAI

Own the cost of a completed task.

ReachMeta

Own the distribution and make switching cheap.

The new model race

Three wedges, one contest

A year ago these three labs were racing up the same leaderboard. This week they are running three different races.

Executive test

Stop asking which model is smartest

Benchmark scores still matter, but this week shows the vendors themselves competing on other terms. Ask the three questions they are optimising for.

01
Work

Can it finish the workflow?

Not answer the prompt: finish the job. Test whether the system completes a whole task across your tools and files, and how much a person still has to fix afterwards.

02
Economics

What does a successful task cost?

Token prices mislead. Count the full cost of a completed task, including reasoning, tool calls, retries and rework. A cheaper price per token can still lose on cost per result.

03
Reach

How fast does it enter real work?

The best model you cannot deploy loses to a decent one already inside your systems. Weigh integrations, compatibility and where your people already spend their day.

Key Takeaways

The operating strategy now matters as much as the model

Model choice is becoming workload routing, task economics and distribution strategy.

01

The unit of competition is completed work

GPT-5.6's tiers plus ChatGPT Work, GPT-Live and preferred placement in Microsoft 365 Copilot show OpenAI packaging capability around finished outputs and long-running workflows. Judge systems the same way: on what they complete, not what they can answer.

02

Cost per task beats cost per token

Grok 4.5's strongest signal is near-frontier intelligence at $0.31 per measured benchmark task, from lower prices and fewer tokens together. Reliability is a separate line item: its hallucination rate roughly doubled even as accuracy improved.

03

Distribution can turn a comeback into a threat

Muse Spark 1.1 does not need to lead a single benchmark to matter. A compatible API at a quarter of rival prices, attached to Meta's reach, changes what everyone else can charge.

FAQ

Yes. GPT-Live and Grok 4.5 arrived on 8 July, then GPT-5.6 general availability, ChatGPT Work and Muse Spark 1.1 on 9 July. GPT-5.6 had been in a limited preview since 26 June.