Decision report

The consumer is unmeasured

Twelve independent research threads, run today, on what the AI expert-data companies actually sell, whether it lasts, and where the durable value sits in health AI.

Prompted by a Forbes item circulating on 7 September 2026: AfterQuery, reported at $3.2bn, selling human professional reasoning to AI labs. Every figure below carries its source. Two of my own earlier conclusions were withdrawn during the day and both are marked.

The answer, in four facts

Every measurement and safety system in AI was built for the professional. The ordinary person is unmeasured, ungoverned and unrepresented, while being roughly 300 million people a week. Four independent registers return zero.

01

Health appears in no frontier lab's tracked risk list

Verified at source in all three safety frameworks.

OpenAI tracks biological and chemical risk, cybersecurity and AI self-improvement. Anthropic tracks two chemical-biological thresholds and two autonomy ones. Google tracks chemical, biological, radiological and nuclear risk, cybersecurity, machine-learning research and harmful manipulation. Google's most recent safety report contains zero occurrences of health, medical, clinical, patient, mental health, self-harm, suicide or sycophancy across the whole document. The UK and US national safety institutes mirror the labs. Nobody independent evaluates clinical advice anywhere.

02

When a lab measures a layperson, it measures whether the model harms them, never whether it helps them

Counted from seven primary model cards.

In eight of ten professional domains, no consumer-facing benchmark exists in any lab's reporting at all. Claude Opus 5's card carries 43 expert-framed capability evaluations and zero lay-framed ones. Every lay-facing test that does exist in any card is a safety, bias or refusal metric. Medicine is the only domain on earth where a lab publishes both sides in the same table, which is why the gap was visible there first.

03

The safety lives in the wrapper, not in the model

Anthropic's own system card, 1 September 2026.

On the raw developer connection with no system prompt, the model responds appropriately in 60% (±14%) of multi-turn conversations about suicide and self-harm. The same model inside the consumer app scores 94%. Every third-party health app builds on the raw connection. Anthropic says so in the card: it encourages developers building on the interface to apply comparable safeguards. Separately, OpenAI's specification puts the self-harm rule at a level nobody can override, and the medical-advice rule at a level any paying developer can.

04

No regulator has ever acted against AI advice for being wrong

Every action found across three US agencies.

Each one is about a misrepresentation concerning the product: not free, not really AI, not really optimising, not really a lawyer. And the federal posture reversed on 22 December 2025, when the US consumer regulator vacated its own order as an undue burden on AI innovation. The health-breach reporting system has no field in which an AI-caused breach could be recorded, because its technical rules have not been amended since January 2013 and contain no word for a model.

So the feedback loop that makes medicine safe, a denominator, an adverse-event channel and a registry, does not exist for the largest health intervention ever deployed.

What the numbers look like

A generation of capability landed on the clinician's side of the line

OpenAI's own health benchmarks, GPT-5 to GPT-6 Astra, on the length-adjusted scores the labs publish.

Clinician tasks Consumer tasks
70 60 50 40 30 20 46.2 60.5 63.4 34.7 25.4 33.1 36.3 Clinician Consumer GPT-5 GPT-5.1 GPT-5.6 GPT-6 Score out of 100. No clinician figure is published for GPT-5.1, so that series has three points.
Read the caveat with the chart. These are the length-adjusted scores. On the unadjusted column in the same table the consumer line falls instead: 41.6 at GPT-5 against 37.8 at GPT-6, while the clinician line rises 51.0 to 69.5. The adjustment penalises the consumer benchmark 2.7 times harder than the clinician one. Both columns are OpenAI's own, and two of my research threads read them opposite ways.

What a card measures, by whose seat the task is written from

Capability evaluations counted in each published model card.

Written from the professional's seat Written from an ordinary person's seat
Claude Opus 5 43 0 Mythos Preview 15 0 Gemini 3.1 Pro 15 0 Gemini 3.8 Flash 14 0 GPT-5.6 launch 10+ 0 Tests written from a professional's seat From a lay seat
The right-hand column is the finding. Every lay-framed evaluation that appears in any of these cards is a safety, bias or refusal metric, not a capability one.

What an hour of expertise is worth to an AI company

Advertised or reported rates, against what the same hour earns clinically.

A US clinical hour, $168 Expert witness $450–600 Surge, medical fellow $250–450 AfterQuery, doctor $120–200 Mercor, internist $130–180 Mercor, lawyer $70–110 AfterQuery, finance $50 General annotation $25–60 Doximity PeerCheck $0, a byline $0 $300 an hour $600
Every bar is a range, and the dashed line is what an American physician's hour already earns clinically, at $386,000 on a 49-hour week. Only one AI buyer clears it outright. Expert witness work, in grey, still pays far more than any of them. And 10,000 physicians already review AI clinical answers on Doximity's PeerCheck for a byline and no fee, with Eric Topol and a former US Surgeon General co-chairing the board.

How long a new benchmark stays useful

Months from release to saturation, oldest first. The half-life has compressed roughly fourfold in five years.

SuperGLUE20 GPQA Diamond27 SWE-bench Verified25 Humanity's Last Exam20 ARC-AGI-213 METR time horizon11 0 15 months 30
A study of 60 benchmarks found 29 already saturated, and that keeping the test private gave no protective effect at all, which removes the obvious defence for anyone selling evaluations as an asset.

The three businesses, ranked

NOSelling expert reasoning to AI labs

This is what AfterQuery does, and Mercor and Surge at larger scale. It is real: Mercor passed roughly $2bn gross annualised revenue by June 2026, and Surge advertises a board-certified Medical Fellow at $250 to $450 an hour. But every vendor describes project fees rather than subscriptions, not one discloses renewal, retention or churn, and the one hard case is brutal: Scale AI went from about $2bn in 2025 to guidance of just over $1bn for 2026 after losing Google, OpenAI and xAI within days of Meta buying half of it. The reason was structural, not reputational: what you ask a vendor to label reveals what you are building.

It is also a capital-equipment cycle wearing a subscription's clothes. Demand is a derivative of the rate of capability improvement, so it does not decay slowly if progress plateaus, it stops within one purchasing cycle.

NOIndependently auditing AI

I got this wrong I told you this morning that Europe had made independent audit a legal obligation five weeks ago. The primary text says the opposite. Article 43 of the EU AI Act puts most high-risk AI on internal self-assessment "which does not provide for the involvement of a notified body", and Article 55 has frontier model providers evaluate themselves. Where Europe does name independent external evaluators, it offers them free access and a non-retaliation protection, with no fee mentioned anywhere: a bug-bounty shape, not an audit market.

The obligation is receding on three fronts. US banking regulators replaced their model-risk guidance on 17 April 2026 and put generative and agentic AI expressly out of scope. Colorado deleted its duty of care in May 2026. Europe deferred its high-risk obligations to December 2027. Exactly one law in force anywhere requires an outside party to examine a deployed AI system, New York City's hiring-tool rule, whose penalties are $500 to $1,500 and which the city comptroller found is not being enforced. The best hard market figure, from a UK government study: £1.01bn total, but only £0.36bn across 84 genuinely independent firms, while the largest group by count is developers assuring their own products.

THE ONEWatching health AI after it is deployed

Not authoring a standard, which is a one-off cost, and you were right to press on that. Running one, continuously, against a model that changes every few weeks and a population that drifts. The head of the US drug regulator is on record that no American health system can validate an algorithm it has already deployed; Stanford needed about 115 hours of expert time to audit two models; the American Medical Association's own digital health lead put it as "we have no standards".

And the reason this one survives is who pays. Where no regulator compels anyone, the only reliable buyer is the party whose balance sheet is exposed to the outcome. Two threads reached this independently without seeing each other: the assurance study found the single structure that works anywhere is insurance-backed, where the evaluator bears the cost of being wrong, and the security study named the malpractice and cyber carriers as the eventual buyer. Nobody has to buy an audit. A carrier that must price this risk has to buy the measurement.

Cybersecurity, which you asked about separately

The attack works, the casualty does not exist yet, and nobody is standing in the gap. All three halves are load-bearing.

Proven. Sub-visual instructions hidden inside medical images fooled all four frontier models across 297 attacks, invisibly to a human observer. A published attack added and removed lung cancer from CT scans convincingly enough to fool three expert radiologists and the AI, with a covert test run on a live hospital network. A model trained on 378,000 clinical notes reproduced patient timelines verbatim, recovering a sensitive diagnosis at 0.91 and 0.82 accuracy.

Not yet realised. Three authoritative registers agree independently. MITRE's catalogue of real AI attacks holds 72 case studies and contains zero hits for healthcare, hospital, clinical, patient or medical across the entire file. The US exploited-vulnerability catalogue holds 1,695 entries, of which nine are AI and all nine are ordinary web flaws in AI plumbing, with no model-level attack at all. And the industry's own 2026 threat ranking concedes that when ranked on 7,714 real incidents rather than a practitioner vote, prompt injection falls out of the top ten entirely.

Unoccupied. Of roughly 100 organisations backing that threat standard, exactly one is a healthcare company. The best-funded hospital AI platform's entire public site returns zero for attack, threat, adversarial, exfiltration, red team or jailbreak: it sells governance, not security. Meanwhile the AI-security field has already consolidated into the big platforms, with four specialist firms acquired.

One thing nobody has checked, and it is cheap

Whether a single AI-caused healthcare breach has ever been recorded. The US regulator publishes 7,184 archived breach narratives. A search of them would settle it: a hit would be the first documented case in existence, and a clean zero would be a strong negative that nobody else holds. The attempt was hard-blocked by one of your own safety guards, which matched the word "submit" on what was a read-only public government search box. The thread did not reword its way around the guard, which was the right call.

The twelve threads

Each ran independently against its own brief, with instructions not to work from assumption and to tag every figure with its source. Open any one for what it found.

Where the threads disagree

Contested

Did the consumer capability gap close? Two threads read the same OpenAI table and reached opposite conclusions. On the length-adjusted scores the consumer benchmark rose 33.1 to 36.3 at GPT-6, its largest move in the series; the thread that found this recomputed the adjustment by hand and reconciled it to 0.1 points. On the unadjusted numbers in the same table the pattern survives: 41.6 at GPT-5 against 37.8 at GPT-6, while the clinician line rises monotonically. The adjustment penalises the consumer benchmark 2.7 times harder, so the two series are arguably not on a common scale.

Both columns are OpenAI's own. I told you at one point that the gap had closed. Treat it as open.

What nobody could establish

Several of these are more interesting than the findings, and each is a cheap experiment for somebody.

What I got wrong today

Three confident answers, two withdrawn. Both were single-sourced when I gave them, and the findings that survived are the ones several independent threads reached separately.