Why Brain AI
Capability is no longer the constraint. Evidence is.
Most enterprise AI never makes it out of the pilot, and not because the pilots fail. To move a decision into production the output has to be reliable enough to act on, and somebody has to be able to show how it was reached. This page sets out the evidence for both, and where the current approaches stop.
Reliability
Choosing a better model does not settle this
The top four labs now sit within 22 rating points of one another, so model choice is becoming a question of cost, latency, domain fit and reliability rather than a search for the single best one. The reliability part does not resolve by picking differently.
22
Rating points
The spread separating the top four labs, which is why model choice is now an operating decision rather than a search for a winner.
Stanford HAI AI Index 2026.
7 to 18%
Hallucination, grounded
Range across relevant frontier models when summarising supplied documents. Grounding helps considerably. It does not finish the job.
Vectara HHEM-2.3 leaderboard, 11 May 2026.
Below 70%
Overall factuality
Every model in the December 2025 FACTS run, across parametric knowledge, search, multimodal input and grounding.
Google DeepMind FACTS, December 2025.
Artificial Analysis’ AA-Omniscience measures a further dimension: whether a model recognises when it does not know, rather than guessing. Calibration differs materially between models that are otherwise closely matched on capability. These benchmarks measure different tasks and are not directly comparable with one another.
There is no single hallucination rate for a frontier model. Reliability depends on the task, the context supplied, the model’s calibration and the controls around it. That is why trust has to be engineered at system level rather than selected at model level.
The evidence requirement
You now have to show how the decision was reached
The EU AI Act is the most comprehensive framework any government has brought into force, and the parts that reach general purpose AI are already live. The Digital Omnibus deferred the high-risk obligations under Annex III to December 2027. It did not defer the rest.
| EU AI Act | Status | What it covers |
|---|---|---|
| Article 5, prohibited practices | In force since February 2025 | An outright ban on specified uses, unaffected by the deferral. |
| General purpose AI obligations | In force since August 2025 | Transparency, documentation and copyright duties on providers of general purpose models. |
| Article 50, transparency | In force since 2 August 2026 | Disclosure that a person is interacting with an AI system, and labelling of AI-generated content. It applies by system function rather than risk tier, so mainstream generative AI deployments are in scope regardless of the Annex III deferral. |
| Annex III high-risk | Deferred to 2 December 2027 | The obligations most often cited as delayed. The three rows above were not. |
The direction is not confined to Europe. South Korea’s AI Framework Act came into force on 22 January 2026, making it the second jurisdiction with comprehensive AI legislation. China enforces binding rules including algorithm registration and mandatory labelling of synthetic content since September 2025. Japan’s AI Promotion Act took a deliberately non-binding route. Brazil’s AI bill has passed the Senate and sits with the Chamber of Deputies. Seventy-two countries now have AI policies of some kind.
Sector supervisors have moved in parallel
| Instrument | Date | What it requires |
|---|---|---|
| FINRA 2026 Annual Regulatory Oversight Report | Published 9 December 2025 | A standalone section on generative AI for the first time. It defines hallucination as output that is inaccurate yet presented as factual, and sets continuing human monitoring as the expectation. |
| NAIC AI Systems Evaluation Tool | Pilot March to September 2026 | A twelve-state pilot giving regulators a structured way to examine how an insurer uses AI, how it governs it, and which models are high risk. |
| California SB 1120 | In force since 1 January 2025 | Requires that final determinations of medical necessity in utilisation review be made by a licensed physician. AI may not autonomously deny, delay or modify care. |
| New York DFS | In force | Requires insurers to demonstrate their systems do not act as a proxy for protected classes. |
The obligation has moved from having a policy to evidencing that the policy was applied to a specific decision.
What it costs when neither is in place
Fluent, plausible, and acted on
None of the following were caused by a weak model. Each output was fluent, well formatted and plausible enough that a trained professional read it and relied on it. More than 1,600 court decisions worldwide have now addressed reliance on AI-generated material that turned out not to exist.
Source: Damien Charlotin, AI Hallucination Cases database. The database held 1,598 cases in June 2026 and passed 1,668 in July. It counts decisions where a court or tribunal found or implied that a party relied on hallucinated material.
| When | Sector | What happened |
|---|---|---|
| July 2026 | UK government | An Upper Tribunal found the Home Office refused an asylum claim citing a country policy note that never existed. |
| April 2026 | Law | Sullivan and Cromwell apologised to a US bankruptcy court over roughly 28 erroneous citations in an emergency motion. |
| October 2025 | Professional services | Deloitte Australia partially refunded A$440,000 to the Australian government for a report containing fabricated citations. |
| Ongoing | Health insurance | A US federal court allowed discovery into whether an insurer used AI to supplant physician decision-making in care denials. |
Not one of these was stopped before it reached the person it affected.
Where we sit
The labs build the model. We build the decision.
The best model in the world does not know what one enterprise, in one jurisdiction, is permitted to say to one customer. That question is answered per tenant, and somebody has to answer it.
Why will the frontier labs not simply build this themselves?
Because the answer is different for every customer, and none of it is a model problem. Every governed decision we run spends money with the labs. We are their customer, not their competitor.
Whose policy
Frontier labs align their models to a single policy applied to everyone, and they do it well. What no lab can do is encode your policy, for your tenant, in your jurisdiction.
Enforceability
Prompting a model to respect a policy makes compliance likely. CIP makes parts of it structural. An out-of-policy tool call is refused at the runtime rather than discouraged in an instruction, evidence outside the policy scope cannot be retrieved at all, and the finished output routes to a named reviewer.
Independence
A model checking its own output is a trading desk marking its own book. Our evaluator is a separate layer that scores the reasoning trace as well as the answer.
AWS shipped CloudWatch before Datadog existed. Datadog became a large company on top of a platform that already had a first-party product, because depth in one problem beat breadth across all of them.
The alternatives
What the current approaches do, and where they stop
Every category below is credible in its own layer, and most enterprises need several of them. None of them compiles the policy, produces the decision under it, and independently checks both. Gartner named AI governance as a market in June 2026 and the category has already split into distinct groups that are often compared with one another when they should not be.
Governance platforms
What it does well
Policy registries, risk tiering, framework mapping and audit evidence packets.
Where it stops
They sit beside the decision, not in it. The evidence they hold is only as good as what the runtime feeds them, and runtime enforcement is roadmap for several of them.
Runtime guardrails
What it does well
A single control point on every model and agent call. Input and output filtering, access control, tenancy and audit logging generated from real traffic.
Where it stops
They gate the boundary. They have no view of whether the reasoning inside was sound, and no mechanism to turn a caught failure into a standing constraint.
Agent frameworks
What it does well
Strong orchestration, tool use and state handling. Fast to build with, and the labs ship their own.
Where it stops
Policy lives in prompt text. Evaluation, where it exists, is the same model grading itself. Nothing is compiled and nothing closes the loop.
Building it in-house
What it does well
Full control, exact fit, and no vendor in the decision path.
Where it stops
Twelve to eighteen months of platform work before the first decision ships, then permanent maintenance, and the evidence has to be rebuilt for every audit.
Category framing informed by the Gartner Magic Quadrant for AI Governance Platforms, published 16 June 2026.
Brain AI compiles the policy, produces the decision under it, and has a separate layer check both.
Human authorisation
The platform prepares and proves the work. A person decides.
Brain AI prepares and proves the work. The clinician, the complaints handler, the money laundering reporting officer decides and approves. Every deployment routes output to a named reviewer, and nothing reaches the person it affects until that reviewer authorises it.