Commercial read · 8 September 2026
The same twelve research threads, answered as a business question rather than an evidence question. Four markets sized, seven parties and what each will actually pay for, six gaps stated as something somebody feels, and what the next eighteen months do to the timing.
The buyer is the insurer, the product is evidence, and the window is open because the regulation is receding while the deployment is not.
Every other party in this chain either has no money in it, is already served, does not need it, or is not asking. The insurer is the only one with a balance sheet exposed to the outcome, and nobody is selling to it in health.
Four businesses sit in this space. Three are closed and one is open.
Selling expert reasoning to AI labs
Industry-wide gross spend: low single-digit billions, heading to high single digits, roughly doubling a year. This is the AfterQuery business.
It is real money. But every vendor charges project fees, not subscriptions, and not one of them discloses whether a customer comes back. The hard case is brutal: Scale AI went from about $2bn of revenue in 2025 to guidance of just over $1bn for 2026, after Google, OpenAI and xAI all left within days of Meta buying half of it. That was structural, not reputational: what you ask a supplier to label tells the supplier what you are building.
It also does not pay a working doctor. Mercor pays $130 to $180 an hour for an internist against about $168 blended clinically, self-employed, no benefits and no cover. It buys the marginal hour and never the working one. Medicine carries no premium over law on these platforms.
Sources: vendor revenues rebuilt bottom-up (Mercor at roughly $2bn gross annualised by June 2026, up from $760m four months earlier; Scale AI 2026 guidance from its own chief executive); Mercor and Surge published rate bands; US physician income $386,000 on a 49-hour week. All verified at source in threads 01, 02 and 04.
Governance software for hospitals
Occupied eighteen months ago by Qualified Health: $125m Series B in March 2026, led by New Enterprise Associates, Anthropic on the share register, claiming 400,000 users across health systems representing about 5% of US hospital revenue.
Founded by Justin Norden, the Stanford professor you half-remembered. Worth knowing what it is not: its entire public site returns zero for attack, threat, adversarial, exfiltration, red team or jailbreak. It sells governance, which answers whether a model is performing acceptably, not whether somebody is attacking it.
Sources: Fierce Healthcare on the Series B; the company's own site; Stanford faculty profile. Thread 04.
Independent AI audit
The only hard figure that exists anywhere is a UK government study: £1.01bn of "AI assurance" but only £0.36bn across 84 genuinely independent firms, and the largest group by headcount is developers assuring their own products.
I told you on Sunday morning that Europe had made this a legal obligation. That was wrong and I withdrew it: the AI Act puts most high-risk systems on internal self-assessment, and frontier providers evaluate themselves. Exactly one law in force anywhere requires an outside examiner, New York City's hiring-tool rule, with penalties of $500 to $1,500 that the city's own comptroller found are not being enforced. The problem is demand, not supply, and demand is going down.
Sources: EU AI Act Articles 43 and 55 read at source; US banking guidance of 17 April 2026 putting generative AI expressly out of scope; Colorado's repeal in May 2026; UK Department for Science, Innovation and Technology assurance market study. Thread 07.
Watching health AI after it is deployed, sold to whoever carries the loss
Buyer's pool: $9.4bn of US doctor-liability premium written by specialist insurers in 2025, plus a share of a $15.6bn to $16.6bn global cyber-insurance market where North America is about two thirds and healthcare pays 42% above the cross-industry median and grows fastest.
Not authoring a standard, which is the one-off cost you correctly attacked. Running one continuously, against a model that changes every few weeks and a population that drifts. The head of the US drug regulator is on record that no American health system can validate an algorithm it has already deployed; Stanford needed about 115 hours of expert time to audit two models; the American Medical Association's own digital health lead put it as "we have no standards".
The reason this one survives is who pays. Where no regulator compels anybody, the only reliable buyer is the party whose balance sheet is exposed to the outcome. Two of the twelve threads reached that independently without seeing each other. The structure is already proven in another sector: Armilla underwrites AI warranties at Lloyd's with Chaucer, AXIS Capital, Convex, Swiss Re and Greenlight Re behind it, backed by third-party model evaluation, so the evaluator carries the cost of being wrong. It is cross-sector and finance-leaning. There is no healthcare occupant.
Verified at source: "Direct premiums written by those MPL specialist insurers increased 3.6% to $9.4 billion last year", AM Best market segment report, read 8 September 2026. That is specialist insurers only, not the whole market. Armilla's Lloyd's coverholder status and named partners verified on the company's own site.
Weaker, and flagged as such: the cyber figures come from market trackers rather than a primary filing, and they disagree with each other by about a billion. Treat $15.6bn to $16.6bn, two thirds North American, as an order of magnitude and not a number.
Not verified, because it does not exist: the healthcare-exposed slice of cyber premium is not separately disclosed anywhere, and no published figure exists for what carriers spend on risk engineering as a share of premium. I am not going to manufacture a capture rate from those. Against the doctor-liability pool alone this is a nine-figure annual market in the US, and the precise number is not knowable from public data.
Seven parties. Six of them want something they will not pay you for.
| Who | What they actually want | What they pay for today | Will they pay you |
|---|---|---|---|
| The person asking |
To know whether this is serious and what to do next. Roughly 300 million a week. | Nothing. Among 18 to 29 year olds, 38% use it because they have no doctor or no appointment and 29% because of cost. | No. 77% say they are worried about privacy and 65% of those uploaded their records anyway. The user is never the payer here. |
| The physician |
Time back, and cover. | Their own liability premium. | No, and they will not sell you their hour either. 10,000 doctors already review AI clinical answers on Doximity for a byline and no fee at all. |
| The frontier lab |
Capability on measurable tasks, and permission to keep operating. | Expert data on project terms, and evaluations. | Not for more data. OpenAI's chief scientist names monitoring confidence, compute and whatever is least automatable as the bottlenecks. None of them is data. But their own specification lists keeping their licence to operate as a co-equal design goal, so they do buy protection. |
| The health system |
To answer "is this safe, and can I prove it" to a board and a regulator. | Governance software, already bought. | Yes, out of the security budget. Nobody today can answer which patients' records the AI read last month, and that audit is already owed under a rule last amended in January 2013. |
| The insurer |
To price a risk it cannot currently observe. | Nothing in health AI. There is nothing to buy. | Yes, and it is the only party that must. It cannot underwrite what cannot be reconstructed, and the loss is already going into the book. |
| The regulator |
Nothing is being demanded of it. | Nothing. | No. No regulator anywhere has acted against AI health advice for being wrong. Every action found across three US agencies is about a misrepresentation: not free, not really AI, not really a lawyer. |
| The app vendor |
To get through hospital procurement. | Compliance paperwork. | Yes, to close deals. It is being asked whether its product resists attack and cannot answer. It also builds on the raw developer connection, where the model handles multi-turn conversations about suicide appropriately 60% of the time against 94% inside the consumer app. |
Consumer figures from polling of 1,343 US adults. Lab bottlenecks from Jakub Pachocki's essay of 6 September 2026 and OpenAI's published specification. The 60% against 94% figure is Anthropic's own system card of 1 September 2026, plus or minus 14 points. Threads 03, 04, 06, 08, 11 and 12.
Stated as the moment somebody notices, not as a missing benchmark.
A hospital lawyer is asked what the AI read, and nobody can say.
Which fragments of which records reach the model is decided by a similarity search that does not respect who is allowed to see what. The privacy officer already owes that audit trail under a rule whose technical safeguards have not been amended since January 2013 and contain no word for a model.
An insurer is writing cover on a loss it cannot see.
No denominator, no channel for reporting when it goes wrong, no registry. The federal breach form has a closed five-value list with no field in which an AI-caused breach could be recorded at all. The feedback loop that makes medicine safe does not exist for the largest health intervention ever deployed.
A vendor is asked whether its product resists attack and has no answer to give.
One healthcare-specific test for this exists in the world, published in February 2026 as an academic dataset. It is nobody's product, nobody's purchase requirement and nobody's regulatory expectation.
A doctor cannot get the model to treat them as a doctor. This is your own complaint, and it is real.
Refusal is already dialled by who is asking: loosened for verified security professionals, tightened for users flagged as high risk. There is no equivalent setting for a verified clinician. Nothing AfterQuery sells touches this, because it is a liability decision by whoever ships the model, not a gap in its training.
A worried person is confidently reassured, and nobody is counting.
A Nature Medicine study in May 2026 found ChatGPT Health under-triaged 52% of textbook emergencies, sending diabetic ketoacidosis and impending respiratory failure to a 24 to 48 hour review. When a family member played the symptoms down, the triage shifted with an odds ratio of 11.7. No lab publishes a triage accuracy figure of any kind, and every test scores the answer, while the harm is a decision days later outside the conversation.
A health system cannot validate the thing it has already switched on.
Generative AI is embedded in the dominant medical records system with trade reporting of more than 85% of its customers using it, health systems can now build agents that act across workflows on their own, and patients are asking questions answered from their own records. The defensive layer beneath that does not exist, and of roughly 100 organisations backing the industry's AI security standard, exactly one is a healthcare company.
Four trends, and only one of them changes what you should do this year.
xAI cut 500 general annotators in the same announcement as a tenfold expansion of specialist tutors in science, finance, medicine and safety. Commodity labelling pays $1 to $12 an hour; credentialed specialists $85 to $300 and up. The arbitrage on a working American physician's hour is already gone and the gap widens as clinical rates rise.
A study of 60 benchmarks found 29 already saturated and, importantly, that holding a test privately gave no protective effect at all. The half-life has compressed roughly fourfold in five years. OpenAI targets a fully automated AI researcher by March 2028. So a business whose asset is a standard is on a shortening clock; a business whose asset is running the standard is not.
Europe deferred its high-risk obligations to December 2027, US banking put generative AI out of scope in April 2026, and Colorado repealed its duty of care in May 2026. A business whose demand depends on a mandate therefore has no catalyst for at least fifteen months. A business whose demand depends on somebody's balance sheet has one today. That asymmetry is the whole timing argument.
Attacks are proven: hidden instructions inside medical images fooled all four frontier models across 297 attempts, and a published attack added and removed lung cancer from scans convincingly enough to fool three expert radiologists. Casualties are not: the main catalogue of real AI attacks holds 72 cases and none is healthcare, and the US exploited-vulnerability list has no model-level attack at all. From inside, "the risk is over-sold" and "this is the quiet before the first case" look identical. That argues for building what the insurer needs to price the risk, which pays either way, rather than what only pays after somebody is hurt.
Build the measurement layer for deployed clinical AI, sell the evidence half to health systems and vendors as a procurement gate, and sell the surveillance half to the malpractice and cyber carriers as the input to their pricing. One product, two payers, and the payer that matters does not need a regulator to show up.
Skip the physician-data business you originally asked about. If you ever want it, counsel already gave the only workable shape: de-identified historical cases plus doctors grading an AI's answer rather than writing their own, and never live symptoms from an identifiable person, which is the one part no contract can make safe.
One cheap experiment is still undone and would sharpen all of this. The US health regulator publishes 7,184 archived breach narratives. Searching them settles whether a single AI-caused healthcare breach has ever been recorded: a hit would be the first documented case in existence, and a clean zero is a strong negative nobody else holds. The attempt was blocked by one of our own safety guards matching the word "submit" on a read-only government search box.