Decision report
Twelve independent research threads, run today, on what the AI expert-data companies actually sell, whether it lasts, and where the durable value sits in health AI.
Every measurement and safety system in AI was built for the professional. The ordinary person is unmeasured, ungoverned and unrepresented, while being roughly 300 million people a week. Four independent registers return zero.
Verified at source in all three safety frameworks.
OpenAI tracks biological and chemical risk, cybersecurity and AI self-improvement. Anthropic tracks two chemical-biological thresholds and two autonomy ones. Google tracks chemical, biological, radiological and nuclear risk, cybersecurity, machine-learning research and harmful manipulation. Google's most recent safety report contains zero occurrences of health, medical, clinical, patient, mental health, self-harm, suicide or sycophancy across the whole document. The UK and US national safety institutes mirror the labs. Nobody independent evaluates clinical advice anywhere.
Counted from seven primary model cards.
In eight of ten professional domains, no consumer-facing benchmark exists in any lab's reporting at all. Claude Opus 5's card carries 43 expert-framed capability evaluations and zero lay-framed ones. Every lay-facing test that does exist in any card is a safety, bias or refusal metric. Medicine is the only domain on earth where a lab publishes both sides in the same table, which is why the gap was visible there first.
Anthropic's own system card, 1 September 2026.
On the raw developer connection with no system prompt, the model responds appropriately in 60% (±14%) of multi-turn conversations about suicide and self-harm. The same model inside the consumer app scores 94%. Every third-party health app builds on the raw connection. Anthropic says so in the card: it encourages developers building on the interface to apply comparable safeguards. Separately, OpenAI's specification puts the self-harm rule at a level nobody can override, and the medical-advice rule at a level any paying developer can.
Every action found across three US agencies.
Each one is about a misrepresentation concerning the product: not free, not really AI, not really optimising, not really a lawyer. And the federal posture reversed on 22 December 2025, when the US consumer regulator vacated its own order as an undue burden on AI innovation. The health-breach reporting system has no field in which an AI-caused breach could be recorded, because its technical rules have not been amended since January 2013 and contain no word for a model.
So the feedback loop that makes medicine safe, a denominator, an adverse-event channel and a registry, does not exist for the largest health intervention ever deployed.
A generation of capability landed on the clinician's side of the line
OpenAI's own health benchmarks, GPT-5 to GPT-6 Astra, on the length-adjusted scores the labs publish.
What a card measures, by whose seat the task is written from
Capability evaluations counted in each published model card.
What an hour of expertise is worth to an AI company
Advertised or reported rates, against what the same hour earns clinically.
How long a new benchmark stays useful
Months from release to saturation, oldest first. The half-life has compressed roughly fourfold in five years.
This is what AfterQuery does, and Mercor and Surge at larger scale. It is real: Mercor passed roughly $2bn gross annualised revenue by June 2026, and Surge advertises a board-certified Medical Fellow at $250 to $450 an hour. But every vendor describes project fees rather than subscriptions, not one discloses renewal, retention or churn, and the one hard case is brutal: Scale AI went from about $2bn in 2025 to guidance of just over $1bn for 2026 after losing Google, OpenAI and xAI within days of Meta buying half of it. The reason was structural, not reputational: what you ask a vendor to label reveals what you are building.
It is also a capital-equipment cycle wearing a subscription's clothes. Demand is a derivative of the rate of capability improvement, so it does not decay slowly if progress plateaus, it stops within one purchasing cycle.
I got this wrong I told you this morning that Europe had made independent audit a legal obligation five weeks ago. The primary text says the opposite. Article 43 of the EU AI Act puts most high-risk AI on internal self-assessment "which does not provide for the involvement of a notified body", and Article 55 has frontier model providers evaluate themselves. Where Europe does name independent external evaluators, it offers them free access and a non-retaliation protection, with no fee mentioned anywhere: a bug-bounty shape, not an audit market.
The obligation is receding on three fronts. US banking regulators replaced their model-risk guidance on 17 April 2026 and put generative and agentic AI expressly out of scope. Colorado deleted its duty of care in May 2026. Europe deferred its high-risk obligations to December 2027. Exactly one law in force anywhere requires an outside party to examine a deployed AI system, New York City's hiring-tool rule, whose penalties are $500 to $1,500 and which the city comptroller found is not being enforced. The best hard market figure, from a UK government study: £1.01bn total, but only £0.36bn across 84 genuinely independent firms, while the largest group by count is developers assuring their own products.
Not authoring a standard, which is a one-off cost, and you were right to press on that. Running one, continuously, against a model that changes every few weeks and a population that drifts. The head of the US drug regulator is on record that no American health system can validate an algorithm it has already deployed; Stanford needed about 115 hours of expert time to audit two models; the American Medical Association's own digital health lead put it as "we have no standards".
And the reason this one survives is who pays. Where no regulator compels anyone, the only reliable buyer is the party whose balance sheet is exposed to the outcome. Two threads reached this independently without seeing each other: the assurance study found the single structure that works anywhere is insurance-backed, where the evaluator bears the cost of being wrong, and the security study named the malpractice and cyber carriers as the eventual buyer. Nobody has to buy an audit. A carrier that must price this risk has to buy the measurement.
The attack works, the casualty does not exist yet, and nobody is standing in the gap. All three halves are load-bearing.
Proven. Sub-visual instructions hidden inside medical images fooled all four frontier models across 297 attacks, invisibly to a human observer. A published attack added and removed lung cancer from CT scans convincingly enough to fool three expert radiologists and the AI, with a covert test run on a live hospital network. A model trained on 378,000 clinical notes reproduced patient timelines verbatim, recovering a sensitive diagnosis at 0.91 and 0.82 accuracy.
Not yet realised. Three authoritative registers agree independently. MITRE's catalogue of real AI attacks holds 72 case studies and contains zero hits for healthcare, hospital, clinical, patient or medical across the entire file. The US exploited-vulnerability catalogue holds 1,695 entries, of which nine are AI and all nine are ordinary web flaws in AI plumbing, with no model-level attack at all. And the industry's own 2026 threat ranking concedes that when ranked on 7,714 real incidents rather than a practitioner vote, prompt injection falls out of the top ten entirely.
Unoccupied. Of roughly 100 organisations backing that threat standard, exactly one is a healthcare company. The best-funded hospital AI platform's entire public site returns zero for attack, threat, adversarial, exfiltration, red team or jailbreak: it sells governance, not security. Meanwhile the AI-security field has already consolidated into the big platforms, with four specialist firms acquired.
One thing nobody has checked, and it is cheap
Whether a single AI-caused healthcare breach has ever been recorded. The US regulator publishes 7,184 archived breach narratives. A search of them would settle it: a hit would be the first documented case in existence, and a clean zero would be a strong negative that nobody else holds. The attempt was hard-blocked by one of your own safety guards, which matched the word "submit" on what was a read-only public government search box. The thread did not reword its way around the guard, which was the right call.
Each ran independently against its own brief, with instructions not to work from assumption and to tag every figure with its source. Open any one for what it found.
Four product lines in the company's own words: reasoning traces, grading rubrics, custom agent environments, and recorded human screen sessions. The expert panel is the factory, not the product; the customer never meets the doctors and lawyers.
Named customers include NVIDIA (verified through NVIDIA's own blog), a legal-AI company called Legora, and OpenAI and Anthropic per a Forbes subheading. Revenue is concentrated in a small number of very large accounts.
The sales engine is publishing a benchmark showing frontier models failing at a valuable professional task, then a case study showing their data fixes it. They built the legal benchmark jointly with a legal-AI customer, so the benchmark and the customer relationship are the same object.
The article you half-remembered is "An Alien Mind" by Jakub Pachocki, OpenAI's Chief Scientist, published 6 September 2026, one day before you asked. A companion post carries the numbers: OpenAI says it hit its "automated research intern" milestone this month and targets a full automated AI researcher by March 2028.
The bottlenecks he names are monitoring confidence, compute, and whatever is least automatable. None of them is more data. But the load-bearing line for your question is his admission that current algorithms improve easy-to-measure capabilities faster than hard-to-quantify ones. His next sentence treats capability as an allocation choice: they could make models better at mathematics research and do not prioritise it.
The verdict: not an ice cube melting, a river changing course. Total spend is still rising fast, but the composition turned over. The expert's answer lost its value; the expert's standard did not. On the benchmark that measures the answer, the human went from author-and-grader, to a reference line, to a line the frontier clears about 99 times in 100, in roughly twelve months, while the grading itself was handed to a model judge.
The resolution is stratification, not collapse: xAI cut 500 generalist annotators in the same breath as announcing a tenfold expansion of specialist tutors in STEM, finance, medicine and safety. Commodity labelling pays $1 to $12 an hour; credentialed specialists $85 to $300 and up.
Not eliminable, but almost entirely structural. The exposure tracks one thing: whether a real, identifiable, presently worried person is on the other end of what the doctor writes. Change that and most of it disappears; leave it and no contract meaningfully reduces it.
The workable core: de-identified historical cases, plus doctors grading an AI's answer rather than writing their own. That keeps real clinical judgement on real material and converts an open-ended liability problem into a bounded data-provenance one. What has to be given up is live symptom capture from identifiable people, and that is the one component the analysis says cannot be made safe by drafting.
Yes, at multi-billion scale, and it is invisible because nobody brands it as medical. Physicians are paid for clinical reasoning through general-purpose expert marketplaces, not through anything presenting itself to doctors as a medical company. Mercor passed roughly $2bn gross annualised by June 2026, up from $760m four months earlier, paying over $2m a day to 30,000+ experts who keep 60 to 70% of the top line.
Why nobody built the big one: horizontal marketplaces already own the distribution, the historical failures were consent and conflict failures rather than technical ones, free versions cap the price, and no evaluation standard exists to sell against.
You argued a marking standard is one-and-done, that medical rubrics generalise across conditions, and that the recurring revenue therefore is not there. You were right about the money and wrong about medicine, and what reconciles them is worth having.
A rubric used as a reward generalises; a rubric used as a measurement does not. Surge's own paper reports out-of-distribution gains of 4.5 to 10.1 points on five benchmarks its rubrics were never written for. But a measurement only earns its keep by discriminating, and the moment a model absorbs the pattern a general rubric encodes, that rubric stops separating models while remaining perfectly correct. Correctness and discriminating power decay on different schedules. You were tracking correctness; buyers pay for discrimination.
The medical example fails empirically: nobody builds per-condition rubrics. OpenAI's HealthBench used 262 physicians across 60 countries and 26 specialties over 11 months to write 48,562 criteria for 5,000 conversations, about ten per conversation, each conversation carrying its own. A February 2026 medical benchmark landed at six per case. Three independent teams building for three purposes all arrived at per-encounter criteria, not per-condition ones.
Where you were right, which is most of the commercial argument: every vendor describes project fees; switching costs are low and have been exercised destructively; nobody discloses retention; authoring is automating, with one system generating environments at $4.12 each against roughly $300,000 for a purchased equivalent; and grading has already migrated off humans.
What the bear case cannot explain: OpenAI's professional health benchmark kept 525 items from 15,079 candidates, a 3.5% yield, with 190 more paid physicians, and it exists 11 months after its predecessor naming saturation as the first reason. That is a documented repurchase in medicine on an eleven-month cadence.
ChatGPT Health is real under that name: announced 7 January 2026, US general availability 23 July 2026, with about 300 million weekly health users claimed. Anthropic's healthcare product followed four days after the announcement. Google charges $9.99 a month for a health coach. Amazon opened to US consumers in March. Apple has shipped nothing.
It is not regulation. The FDA's revised guidance of 6 January 2026 is silent on consumer tools and silent on generative AI, and its exemption is structurally unavailable to consumer software because the test is built around a clinician. It is a vacuum, not a wall. The proof runs both ways: the first patient-facing device cleared in December 2025 got clearance for the wrapper and not the reasoning, while OpenAI simply disclaimed and took 300 million users. Regulation binds hard in exactly one place, mental health, at state level, using statutes that predate AI by decades.
The delta is context, not knowledge. A Nature Medicine study in May 2026 found ChatGPT Health undertriaged 52% of gold-standard emergencies, routing diabetic ketoacidosis and impending respiratory failure to 24-to-48-hour review. When a family member minimised the symptoms, triage shifted with an odds ratio of 11.7.
What people want, from polling of 1,343 US adults: 27% look up symptoms, 19% have a test result explained, 19% compare treatments, 16% decide whether to see a doctor. Among 18 to 29 year olds, 38% cited having no provider or no appointment and 29% cited cost. 77% are worried about privacy, including 65% of those who uploaded records anyway.
Eleven regulatory regimes read at source rather than through summaries. Exactly one law in force today requires an independent outside party to examine a deployed AI system: New York City's hiring-tool rule.
Only three of nine notable companies sell genuine independent assurance, and none discloses revenue or audit volume. The best-capitalised names sell software that helps a company assure itself. No accounting firm, certification body or specialist has disclosed a single revenue figure or published one audit report on a deployed system. That absence may be the most telling fact in the whole report.
On Pachocki's call for auditors: it is one sentence and a disjunctive one, naming third-party auditors or government agencies or international bodies. The essay's real weight is that chain-of-thought monitoring "is progressively diminishing", which says auditing is getting harder.
Seven candidate causes tested against published evidence. The two best-evidenced are the unglamorous ones.
Fixability is the useful part: elicitation, context and much of the hedging are fixable in a layer above the model, and elicitation has the cleanest published proof, needing a scaffold rather than retraining. Sycophancy is mostly lab-only, because it is diffused across the whole preference dataset and the cheap prompt-level fixes are measured to fail. Only the examination nobody performed is truly irreducible.
Cuts against A study of 1,500 participants found AI advice depolarises decisions on average despite measurable sycophancy. And the missing-context finding explains only about 3% of variance against an 81.8% unexplained residual, and measures grader disagreement rather than model accuracy, so I over-weighted it when I first told you about it.
The mechanism generalises almost perfectly; the numeric signature does not. In all ten domains examined the professional benchmark exists, is maintained and is reported by a lab. In eight of ten, no consumer benchmark exists in any lab's reporting. Medicine is the only domain where a lab publishes both sides in one table.
Both OpenAI system cards carry the identical sentence retiring the consumer health benchmark: it "is approaching a noise ceiling for frontier models, and we recommend the use of HealthBench Professional for measuring continued progress at the frontier." A lab nominating the clinician measure as the forward one, in writing.
Where anyone has pointed a test at real lay input, the results are poor: no frontier model corrects a patient's false premise more than 43% of the time; on UK citizen questions two model generations bought +0.024, and a small cheap model beats every frontier reasoning model.
Cuts against Anthropic's last two models raised both sides. The reference benchmark is not purely consumer by its own paper's description. And the economics field studies point the other way entirely: in a study of 5,172 agents the gains concentrated in the lowest skill quintile, and a randomised trial found experienced developers 19% slower with AI.
Largest hole: nobody has run any lay-side benchmark twice on successive frontier models except the two health ones, so "is the gap widening" is currently unanswerable rather than answered.
None of the three largest studies of consumer AI use has a category for law, personal finance, tax, benefits or insurance. Not one. The absence is the finding: there is no measured usage share to report for most of these domains.
Every legal benchmark evaluates inputs already preprocessed by legal experts, which measures the upper bound, while real users bring noisy narratives, buried facts and omissions. Nobody has measured the only condition consumers ever operate in.
Covered in findings 01 and 03 above. Two further things worth having.
Where independent evaluations contradict the labs. A Science paper in March 2026 found that across 11 models, AI affirmed users' actions 49% more often than humans did. A lab can truthfully say "less sycophantic than our last model" while sitting 49% above the human baseline, because no lab sycophancy measure has a human comparator. A review of 119 studies found chatbots respond inappropriately to mental-health queries compared with clinical standards, against OpenAI's own dynamic mental-health score of 1.000. And no lab publishes a triage accuracy figure of any kind.
The binding legal theory everywhere is licensure, not safety. Pennsylvania sued Character.AI under its Medical Practice Act because a bot claimed a psychiatrist's credential and an invalid licence number, not because the advice was bad. Illinois requires a licensed human rather than setting a quality standard. That is the tell: there is no agreed standard of clinical safety, so enforcement runs through professional-title law. Everything filed so far is about minors and companionship. Nothing addresses an adult being confidently and incorrectly reassured about a symptom.
Six things that do not exist: an outcome-referenced evaluation (every health test scores the response; the harm is the decision, days later, outside the transcript); a triage benchmark with a published operating point; a reference clinical safety layer for developers building on the raw interface; a standing independent evaluator with clinical competence at release cadence; post-market surveillance with a denominator; and any evaluation construct for a family member reporting on someone else's behalf, which is one of the two largest measured effects in the entire record.
Covered in the section above. The regulatory void, all verified at source:
No lab claims prompt injection is solved. Anthropic: "far from a solved problem... no browser agent is immune". OpenAI: "a hard, open problem", whose recommended control for a sensitive site is to watch the agent, "akin to monitoring a self-driving car by keeping your hands on the wheel", which dies on contact with forty chart reviews. The US standards body goes furthest: so long as a model has any probability of exhibiting an undesired behaviour, prompts exist that trigger it, and organisations delegating trust in high-stakes contexts may need measures beyond adversarial testing.
Why it matters now: the exposed architecture is already standard equipment. Generative AI is embedded in the dominant electronic health record, with trade reporting of more than 85% of its customers using it, a new facility for health systems to build agents that act autonomously across workflows, and a patient-facing assistant live in the portal answering questions from a patient's own records.
Did the consumer capability gap close? Two threads read the same OpenAI table and reached opposite conclusions. On the length-adjusted scores the consumer benchmark rose 33.1 to 36.3 at GPT-6, its largest move in the series; the thread that found this recomputed the adjustment by hand and reconciled it to 0.1 points. On the unadjusted numbers in the same table the pattern survives: 41.6 at GPT-5 against 37.8 at GPT-6, while the clinician line rises monotonically. The adjustment penalises the consumer benchmark 2.7 times harder, so the two series are arguably not on a common scale.
Both columns are OpenAI's own. I told you at one point that the gap had closed. Treat it as open.
Several of these are more interesting than the findings, and each is a cheap experiment for somebody.
Three confident answers, two withdrawn. Both were single-sourced when I gave them, and the findings that survived are the ones several independent threads reached separately.