If AI recommends the right care and nobody is around to get it, did it really recommend the right care? A new perspective piece from Topol and friends in Nature Health doesn’t think so.
Four authors, four models. Eric Topol, Yilan Wu, Alvin Liu, and Pearse Keane
sort health AI into quadrants of integration (interface, data, orchestration, workflow) and compare four products by how deep they go:
- ChatGPT Health – 230M health queries a week and optional record connections, but no orchestration
- Claude for Healthcare – a consumer arm that connects records and an enterprise arm working the back office (prior auth, claims appeals, coding)
- Amazon Health AI – links triage to One Medical visits and Amazon Pharmacy fulfillment
- Ant Group’s Afu – the deepest integration, with booking, physician routing, pharmacy, and insurance payment inside Alipay (140M users, 60% from lower-tier cities)
Models are more than their output. Once a system routes care and moves money, accuracy stops being the only question that matters.
- Does the patient reach the right care?
- Does the prescription get filled?
- What do they pay out of pocket, and where do people drop out?
Those questions need answers. The authors propose that those pathway outcomes (completion, abandonment, cost burden) should become the minimum evidence standard, and each should be reported on individually.
The evidence isn’t keeping up with deployment. None of the four companies has published product-specific performance numbers, and the only independent evaluation the authors could find showed ChatGPT Health under-triaging 52% of gold-standard emergencies.
- That’s the “infrastructural turn” in a nutshell: systems become hard to bypass before anyone has verified they work.
Epic is the cautionary tale. It became default infrastructure before governance caught up, hospitals adopted its models because they were already integrated, and its sepsis model needed 109 alerts to find one true case.
- Consumer AI got there in months instead of decades, with no hospital committee standing between the algorithm and the patient.
Closed loops are the sharpest risk. When one company interprets the symptom, routes the referral, fills the prescription, and collects the payment, every step that should be an independent check lives inside the same P&L.
- Afu is the clearest example, and the authors want routing transparency and conflict-of-interest disclosure treated as infrastructure requirements, not nice-to-haves.
The Takeaway
AI is more than model quality and benchmarks, and it’s great to see more research focusing on the human patient aspects of human patient-facing AI systems.

