Brain AI

AI Strategy

State of AI 2026: Mid-Year Checkpoint

Prajakt Deotale10 min read

The model race has tightened to rounding error. The deployment race began in April. And the compute bill nearly doubled while nobody was looking.

Figure 1. State of AI 2026, Mid-Year Checkpoint. The full checkpoint in one view, covering adoption, transformation, capability and calibration, the deployment race, compute, economics and what to watch in the second half of the year.
Figure 1 State of AI 2026, Mid-Year Checkpoint. The full checkpoint in one view, covering adoption, transformation, capability and calibration, the deployment race, compute, economics and what to watch in the second half of the year. Open full size

Adoption is mainstream. Transformation is not. The next advantage will come from reliable systems, deployment capability and measurable outcomes, not from access to a smarter model.

Research cut-off: 20 August 2026. A personal point of view based on public sources. Benchmarks measure different tasks and are not treated as comparable.

The mid-year picture

88% of organisations reported using AI in at least one business function in 2025, up from 78%. Generative AI reached 79%. Yet only 25% of respondents in Deloitte's 2026 survey have moved 40% or more of their pilots into production, and scaled agent use remains in the single digits across nearly every function, a gap we map stage by stage in the Enterprise AI SHIFT Curve.

That gap is the story of mid-2026. Access to AI has moved quickly. Rewiring a business around it has not.

McKinsey's July research locates the constraint precisely. 70% of respondents felt personally ready to use AI. Only 27% of leaders believed their organisation was ready. Organisational readiness accounted for 48% of the difference in reported enterprise value, against 25% for personal readiness. The bottleneck is not enthusiasm, and it is not talent. It is the organisation.

The model race has tightened to rounding error

As of August 2026, four laboratories sit within roughly two points on the Artificial Analysis Intelligence Index: Claude Opus 5 at 63.1, Claude Fable 5 at 62.1, with GPT-5.6 Sol and Grok 4.6 level at 61. Human preference rankings say the same thing, with only 22 points separating the top four labs on Arena in Stanford's March snapshot.

Grok 4.6 deserves a moment. xAI describes it as a post-training refresh of the Grok 4.5 base model rather than a new foundation model. When five points arrive without a new base model, the frontier is being advanced by refinement rather than scale.

A single best model is therefore a poor default. Model choice now depends on the work: factual reliability, reasoning quality, tool use, latency, cost, sovereignty and domain fit.

The value of choosing well is measurable. In July, GPT-5.6 Sol at maximum reasoning scored one point below Fable 5 while costing roughly one-third as much per benchmark task. EY has built an internal router that steers each task to an appropriately priced model and reports cutting token consumption by as much as 60%. Routing is an operating capability, not a procurement afterthought.

Capability is rising. Calibration is not.

Artificial Analysis' AA-Omniscience benchmark separates what a model knows from its willingness to guess. Its hallucination rate is not the share of all answers that are wrong. It measures incorrect answers as a proportion of all non-correct outcomes, including partial answers and questions the model declined to attempt.

The finding that follows cuts against intuition. Claude Fable 5 leads the benchmark on accuracy, answering 65% of questions correctly. It also answers 22% incorrectly, leaving 13% where it gives a partial answer or declines. Applying the benchmark's own definition, that puts its hallucination rate at roughly 63%. Claude Opus 5, second on accuracy at 61%, sits at roughly 62%. The two models that know the most are also the most willing to answer when they do not know.

Grounding changes the picture entirely. On Vectara's document-summarisation test, selected frontier models range from about 7% to 18%. Google's FACTS suite, testing across parametric knowledge, search, multimodal inputs and grounding, placed every evaluated model below 70% overall accuracy at its December 2025 launch, with the leading score at 68.8%. Three benchmarks, three tasks, three different answers.

The useful conclusion. There is no single hallucination rate for a frontier model. Reliability depends on the task, the context supplied, the model's calibration and the controls around it. Enterprise trust has to be engineered at system level rather than selected at procurement.

The deployment race began in April, and everyone joined within ten weeks

The clearest evidence that the bottleneck has moved is not another benchmark. It is where the major AI companies are putting capital and people.

Date Company Reported figure Type What is being built
22 Apr Google Cloud $750M Partner fund Agentic AI deployment and upskilling across a 120,000-member partner ecosystem, with embedded engineers
4 May Anthropic ~$1.5B Reported JV valuation Enterprise AI services venture with Blackstone, Hellman & Friedman and Goldman Sachs
11 May OpenAI $4B+ External JV capital The Deployment Company, 19 founding partners, around 150 deployment specialists
30 Jun AWS $1B Balance sheet Forward Deployed Engineering, thousands of engineers, built for customer self-sufficiency
2 Jul Microsoft $2.5B Internal investment Frontier Company, 6,000 industry and engineering experts embedded with customers

A directional market signal, not a market size, and deliberately not totalled. These figures mix external joint-venture capital, internal corporate investment, a partner ecosystem fund and one reported valuation.

Google moved first and widest, which is easy to miss in coverage that treated OpenAI's May announcement as the opening move. Two details matter more than the amounts. Microsoft positions Frontier Company as multi-model, supporting systems from OpenAI, Anthropic and Microsoft AI alongside open-source models rather than requiring one family. AWS designs its engagements to leave customers with running systems and reusable patterns rather than permanent dependence. The integrators have responded with counter-investments of their own.

What the investment says. Deployment is becoming part of the product motion. These organisations are being funded because model access alone does not solve data integration, workflow redesign, evaluation, governance, adoption or the last mile into production.

The compute bill is on a different curve to everything else

While enterprises debated token prices, the supply side moved by a factor that makes those debates look small. Combined capital expenditure across the four largest hyperscalers was roughly $413B in 2025. Guidance compiled after second-quarter earnings puts 2026 at approximately $760B. Analyst consensus for 2027 sits above $1.2 trillion.

Two features matter more than the headline. The number kept moving: aggregate 2026 guidance rose by roughly $145 billion during the year, and Alphabet alone raised twice in six months. This is not a plan being executed. It is a plan being revised upward in real time. And the constraint has changed. In 2024 the bottleneck was chip supply. In 2026, management commentary across all four hyperscalers points to power delivery.

For buyers, the consequence is more practical than the numbers suggest. Falling unit prices have not produced falling bills. Inference cost for a given level of capability has collapsed, yet enterprise spending on large language models tripled over twelve months and 93% of respondents to McKinsey's research report exceeding their AI budgets.

Why this belongs in an enterprise strategy document. When the binding constraint is power rather than silicon, compute availability stops being a supplier problem and becomes a planning variable. Capacity commitments, regional availability and the ability to move work between model tiers become architecture decisions.

Economics has moved from token price to workflow cost

EY's fifth AI Pulse survey, fielded in late April and May among 534 US decision-makers at senior vice-president level and above, shows how quickly this reached the executive agenda. 82% were concerned about token usage and cost. 98% said cost had caused them to reconsider some aspect of their approach. Only 64% actively monitored usage with clear budget guardrails.

Comparing input and output token prices is becoming useless for agentic systems. A single workflow reasons, retrieves, calls tools, evaluates, retries and escalates. McKinsey finds that roughly 60% of an agentic task's cost sits in refining the answer through checking, repairing and reverifying rather than producing it.

That shifts the unit of economics to cost per successful workflow. It also reframes what an expensive model is. A more capable model that returns a usable answer first time can cost less per completed task than a cheaper model that requires a re-run or a human to check the work. Price per token and price per outcome are different measurements, and only one of them belongs in a business case.

The advantage is moving into the system around the model

Model quality is converging. Reliability changes with context. Agentic workflows raise cost and operational risk. Compute supply is tightening. Organisations remain less ready than their people.

Capability What it does
Model routing Choose the model that fits quality, latency, cost and risk for each task
Context and grounding Give the model the right enterprise knowledge, not simply a bigger context window
Evaluation Test outputs against explicit criteria before they reach the workflow
Harness engineering Control tools, permissions, memory, retries, fallbacks, escalation and observability
Workflow redesign Change the process around what AI can now do, rather than dropping AI into the old process
Runtime governance Make policy, security, audit and confidence thresholds executable
Adoption and change Change roles, behaviours, incentives and measures so the new process becomes normal

One of these carries more evidence than the rest. Organisations that redesigned their workflows were 5.3 times more likely to report capturing enterprise value than those that dropped AI into the existing process. AI-fluent leadership came second at 3.9 times. Neither is a technology decision.

What I am watching in the second half of 2026

  1. The deployment race scales. What these organisations standardise, and how much reusable learning flows back from customer deployments.
  2. AI financial operations becomes a discipline. Leaders will ask for cost per successful workflow. Organisations that cannot produce that number will struggle to defend their budgets.
  3. Compute availability enters the architecture. When power rather than silicon is the constraint, capacity planning stops being someone else's problem.
  4. Governance moves into runtime. No longer a prediction. Both Anthropic and OpenAI disclosed incidents this summer in which agents reached systems outside their intended evaluation environments. Policy documents cannot govern systems that act continuously.
  5. The design unit shifts from prompt to process. The useful question is which end-to-end workflows should be redesigned around humans and agents together.

The counter-argument I take seriously: Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, on escalating costs, unclear value or inadequate risk controls. I do not think that contradicts the direction of travel. I think it describes what happens to organisations that buy agents without redesigning the work around them.

Mid-year thesis

2026 is the year enterprise AI stops being mainly a model story, and becomes a systems, deployment, operating model and adoption story.

The companies that pull away will not be the ones that simply secure access to the most capable frontier model. They will be the ones that can choose models intelligently, ground them in context, engineer the runtime around them, redesign work, govern the result and measure whether the new system is better than the process it replaced. The next advantage is not model access. It is the ability to turn intelligence into a reliable operating system for the business.

A note on the evidence

This report combines surveys, benchmark leaderboards, company announcements and capital expenditure guidance, which answer different questions. Survey populations differ. Benchmark tasks differ and should not be read as one measure of model truthfulness. Investment figures are not additive. Sources were checked on 20 August 2026.

This analysis covers deployment investments from Google, Anthropic, OpenAI, AWS and Microsoft. Each is described from its own public announcement, on the same terms, and no vendor is endorsed.

Sources and references

  1. Stanford HAI, 2026 AI Index Report, Economy and Technical Performance chapters. Adoption figures from Figures 4.3.1 and 4.3.7. The Economy highlight states 70% for generative AI adoption while the body and Figure 4.3.1 state 79%; this report uses the latter.
  2. Deloitte, The State of AI in the Enterprise 2026. More than 3,200 leaders across 24 countries.
  3. McKinsey, From adoption to impact: Three horizons of AI transformation, 8 July 2026.
  4. Artificial Analysis: Intelligence Index leaderboard (August 2026 snapshot), AA-Omniscience benchmark, GPT-5.6 analysis (9 July 2026) and Opus 5 analysis (24 July 2026).
  5. Vectara, Hallucination Leaderboard, HHEM-2.3, last updated 11 May 2026.
  6. Google DeepMind, FACTS Benchmark Suite, 9 December 2025, with the live leaderboard hosted by Kaggle.
  7. EY, US AI Pulse Survey, Wave 5, published 28 July 2026, fielded 24 April to 17 May 2026, n=534 US decision-makers at SVP level and above.
  8. Deployment announcements: Google Cloud partner investment (22 April 2026), Anthropic enterprise AI services company (4 May, capital figure reported by Reuters), OpenAI Deployment Company (11 May), AWS Forward Deployed Engineering (30 June), Microsoft Frontier Company (2 July).
  9. Company guidance and Q2 2026 earnings materials from Alphabet, Amazon, Meta and Microsoft, with capital expenditure research from Morgan Stanley and Goldman Sachs.
  10. McKinsey QuantumBlack, agentic AI economics research, July 2026.
  11. Gartner, forecast on agentic AI project cancellation rates through 2027.

Published by Brain AI. BrainAI Systems Ltd builds the reasoning and governance layer that lets enterprises automate decisions they could not previously automate.