Week of July 11, 2026
OpenAI targets the work. xAI targets the cost curve. Meta targets a comeback.
The signal
Within two days, OpenAI, xAI and Meta each gave a different answer to the same question: how do you win when every lab has a strong model?
OpenAI's move
GPT-5.6 splits OpenAI's flagship into three durable tiers: Sol, Terra and Luna. You choose how much intelligence, and how much cost, each task deserves.
Scores and task costs are Artificial Analysis measurements at each model's highest reasoning effort. The strategy sits around the models: the same week brought ChatGPT Work, the GPT-Live voice model, a unified desktop app and preferred placement in Microsoft 365 Copilot. OpenAI is selling the finished work, not the chat window.
xAI's move
Grok 4.5 comes from xAI, which folded into SpaceX and rebranded SpaceXAI this month. Its headline result is not the score. It is what each completed task costs.
US$ per completed Intelligence Index task
$0.31
per benchmark task
Artificial Analysis measured Grok 4.5 at $0.31 per Intelligence Index task: near-frontier intelligence at the task price of Kimi K2.6, which scores 11 points lower. Price is only half of it. Grok needs about 14,000 output tokens per task, far fewer than its rivals.
Cheap is not the same as reliable
On the same evaluations, Grok 4.5's hallucination rate rose to 54%, from 25% for Grok 4.3, even as accuracy improved. Budget for reliability testing before routing real work to it.
Artificial Analysis Intelligence Index at each model's tested reasoning configuration, retrieved 11 July 2026. Measured benchmark cost, not a production quote.
Meta's move
Muse Spark 1.1 is Meta's first paid, proprietary model. That is a deliberate break from the open-weight Llama strategy that built its developer following.
Meta reports a Terminal-Bench score of 80.0. Independent evaluator Vals AI measured 69.29. Artificial Analysis has not yet published results for Muse Spark 1.1.
Treat the comeback as an attempt, not a result. It can still move your market: a cheap, compatible API attached to Meta's distribution forces rivals to respond.
The new model race
Put the three moves on one map and the week's real story appears: the advantage now sits around the model as much as inside it.
The new model race
Three wedges, one contest
A year ago these three labs were racing up the same leaderboard. This week they are running three different races.
Executive test
Benchmark scores still matter, but this week shows the vendors themselves competing on other terms. Ask the three questions they are optimising for.
Not answer the prompt: finish the job. Test whether the system completes a whole task across your tools and files, and how much a person still has to fix afterwards.
Token prices mislead. Count the full cost of a completed task, including reasoning, tool calls, retries and rework. A cheaper price per token can still lose on cost per result.
The best model you cannot deploy loses to a decent one already inside your systems. Weigh integrations, compatibility and where your people already spend their day.
Key Takeaways
Model choice is becoming workload routing, task economics and distribution strategy.