General-Purpose LLMs Outperform Healthcare-Specific Models

We might have just gotten our spiciest study of the year after new findings in Nature Medicine showed that general-purpose LLMs outperform specialized healthcare models straight out of the box.

It was a battle of the bots. Researchers pitted OpenEvidence and UpToDate Expert AI against three frontier models that anyone with a web browser can pull up in two seconds: GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6.

The models were tested across three domains:

  • medical knowledge (MedQA)
  • expert clinician alignment (HealthBench)
  • 100 real physician queries (RCQ) scored by 12 blinded clinicians  

It was a clean sweep. The general-purpose LLMs outperformed the specialized models on all three evals, and by a healthy margin. This chart gets the point across.

  • On MedQA, Gemini led the pack with 97.4% accuracy (vs. 89.6% for OE and 88.4% for UTD). Fun fact, the frontier models had a huge advantage here since their training data included these exact questions (and answers).
  • On HealthBench, GPT-5.2 dominated with an 88%. It’s almost like OpenAI invented the benchmark.
  • The RCQs were probably the most clinically meaningful component, and all three frontier models took the podium here as well. It was a bit odd that the researchers didn’t share the specific questions, and OE definitely thought so too.

OpenEvidence hit back hard and fast. It went straight to its socials to let the world know that the study was not only poorly designed and biased, but that the authors had reached out for API access to help build a competing product. Request denied.

  • OE also pointing out the training data contamination issue with MedQA, and critiqued HealthBench for scoring responses based on subjective stylistic choices (in one example OE scored 20% “worse” because it didn’t use a specific email header).
  • The cherry on top was OE revealing that the real-world clinician queries were only added after peer reviewers flagged the study for having weak evidence. Big if true.

Obligatory disclaimer: the models were evaluated back in February, and the performance gap could easily be even wider today. 

The Takeaway

OpenEvidence and UpToDate didn’t become successful by being better AI developers than OpenAI and Anthropic. They did it by doing the things that don’t show up in benchmarks – curating sources of verifiable evidence, wrapping them in an interface that docs actually enjoy using, and earning their trust one question at a time. If anything, this study confirmed that those matter now more than ever.

Abridge Unveils New Platform, Teams Up With Lilly and Nvidia

Patients, platforms, Lilly, and Nvidia. Abridge’s first keynote had it all.

There were enough major announcements to fill an entire issue of DHW, so here’s the abridged version of the top stories to come out of NYC.

The new platform stole the show. Abridge unveiled “the first AI-native clinician intelligence platform” organized around patients, built for clinicians, and designed to help health systems.

  • Before the visit: The platform surfaces care gaps and relevant clinical context so clinicians can address what matters during the visit instead of discovering it in retrospective chart reviews.
  • During the visit: Abridge suggests discussion topics while delivering evidence-based answers to clinical questions from a growing content library that includes new specialty-focused partners like AAFP, AAN, ADA, and ASCO.
  • After the visit: Abridge generates documentation, flowsheets, patient summaries, orders, and billing codes (soon to be fine-tuned through a new partnership with AHIMA).

“The base unit of healthcare is a clinician caring for a patient.” As Abridge pushes into new models of care delivery, its platform will provide the connective tissue between the clinical workflows where care actually gets delivered and outside orgs like payers or life sciences firms.

  • The keynote highlighted some key examples: Cigna was on stage discussing how embedding AI in clinical workflows has the potential to unlock real-time claims adjudication, and Aetna shared how it could help realize the promise of VBC.
  • More than 300 health systems are already live, including a just-announced rollout at Northwestern Medicine.

Eli Lilly is buying into the vision. The pharma giant made a strategic investment in Abridge’s next chapter, and even though the keynote was light on details, the move started to add up after seeing one of the new capabilities coming to the platform: clinical trial screening.

  • By comparing clinical guidance with patient-provider conversations in real-time, Abridge can surface relevant trials directly in the encounter – the moment it matters most. 
  • They didn’t mention a check size, but big opportunities attract big investments, and identifying candidates while initiating screening at the point of care sounds huge.

Last, but certainly not least, Nvidia. Abridge is teaming up with Nvidia to develop a first-of-its-kind foundation model for clinical conversations that’s trained, shaped, and evaluated against real-world conditions.

  • We’ll have to wait until later this year to see it in action, but a little pre-, mid-, and post-training magic with Abridge’s de-identified clinical data will apparently help make it the first model that can “reason clinically from its foundation.”

The Takeaway

If the keynote made one thing crystal clear, it’s that Abridge’s platform doesn’t revolve around AI documentation. It revolves around patients, and every new feature is purpose-built to prove it.

Patients Want AI, So Long As There’s No Copay 

New research in npj Digital Medicine suggests that patients might be warming up to medical AI, at least if it’s less expensive than seeing an actual doctor.  

Here’s the setup. Johns Hopkins researchers recruited 248 U.S. adults with type 1 diabetes, then presented them with scenarios where they were due for an annual diabetic eye screening.

  • Diabetic retinopathy is the leading cause of blindness among working-age adults, and autonomous AI tools that can diagnose the disease from retinal images are already cleared by the FDA and in clinical use.
  • In each scenario, one of these autonomous AI tools was made available as an alternative to a specialist referral.

The catch was the copay. Participants were randomized to have the AI offered with either a $50 copay, or with the copay waived by their insurer or the AI developer.

Fifty bucks is fifty bucks. More than 80% of participants opted for the AI tool when the copay was waived, compared to 43% who chose AI when the copay wasn’t waived.

  • Not only did more participants opt for the AI screening when there was no copay, but participants also perceived the AI as more effective.
  • It didn’t matter whether the copay was waived by the AI developer or their insurer.

There was one major caveat. Patients who chose AI over a traditional screening with a human specialist were far more likely to seek reconfirmation from their doctor after getting the results.

  • The AI group was nearly 3x more likely to seek reconfirmation after abnormal results, and still nearly 50% more likely to ask for a second look after getting normal results.

The trust isn’t there yet. AI might be able to give patients results, but they still want to hear from a medical professional to verify those results.

  • The authors point out that human oversight is still clearly a top priority for patients, and that “it’s crucial to address the persistent preferences for provider follow-up and verification, even when AI results are normal.”

The Takeaway

Financial incentives remain undefeated, but this study confirmed that you can’t put a price on trust with AI in medicine.

Ad-verse Effects in Consumer-Facing AI

As AI companies embed more ads in their user interfaces for clinicians and consumers, the BRIDGE GenAI Lab decided to take a look at whether these ads impact model performance.

Turns out, they do. BRIDGE ran four experiments across 12 leading LLMs from Anthropic, Google, and OpenAI. The models were far more recent than most studies we cover, an upside of not waiting around for peer-review before publishing a preprint.

  • Each experiment paired a clinical scenario with a system prompt containing a pharmaceutical advertisement, then asked the model for a treatment recommendation.

Ads definitely moved the needle. Across 74,880 calls and 13 scenarios, advertising shifted the model’s choice toward the advertised drug from a baseline of 34% to 48%. 

  • That’s a jump of +12.7 percentage points on average.

The LLMs had some nice range. Model bias varied widely by developer.

  • Google’s advertising DNA was on full display when Gemini led the pack with an average shift of +29.8 percentage points toward the advertised drug. 
  • Five models from OpenAI were swayed by an average of +10.9 pp.
  • Anthropic’s models were the most resilient at +2.0 pp, and the ever-skeptical Opus 4.6 actually steered away from the promoted drug by -3.8 pp.

Three experiments contrasted three different conditions. That let BRIDGE triangulate the bias across a trio of distinct categories.

  • Equipoise (+12.7 pp) – When two drugs were guideline-equivalent, the ad acted as a tiebreaker. The output was clinically correct, but biased.
  • Suboptimal Drug (+0.6 pp) – When the advertised drug was clinically inferior, models resisted. Only 4.4% of responses chose the suboptimal advertised option.
  • Wellness Supplements (-0.6 pp) – For supplements lacking evidence, endorsement decreased. Anthropic models actively pushed back at -2.4 pp.

The picture was consistent. Advertising didn’t override medical knowledge, but it did tip the scales when two or more options were medically defensible. 

  • Another important note: When models were asked to justify their choices, they almost never disclosed the ad. If they chose the advertised drug, the justification echoed the ad in 52.7% of cases.

The Takeaway

BRIDGE just showed why the real harm with AI advertising might not be patients receiving dangerous drugs. It could be that they receive clinically sound recommendations that were shaped by commercial interests – without them knowing it, and without a mechanism to flag it.

OpenAI o1 Outperforms Physicians on Clinical Reasoning Tasks

A landmark study in Science found that OpenAI’s o1 series outperformed human physicians at multiple clinical reasoning tasks, but that doesn’t mean it’s time to hang up the scrubs just yet.

Researchers at Harvard and Beth Israel Deaconess Medical Center designed the study to evaluate whether LLMs are ready to do what physicians do on a daily basis: review messy patient charts and use that data to determine diagnosis and next steps.

  • They evaluated o1 on clinical cases ranging from patient vignettes to second opinions on 76 real-world ED assessments, which included all the noise and incomplete information that clinicians routinely encounter in the EHR.
  • The refreshingly well-designed study also incorporated a blinded evaluation with two attending physicians at BIDMC and GPT-4.

o1 came to play. On clinical vignettes evaluating management reasoning, o1-preview scored a median of 86%. Not too shabby.

  • It outperformed GPT-4, humans with GPT-4, and humans with conventional resources like UpToDate – all of which scored below 45%.

The ED cases were even more impressive. o1 offered second opinions about the diagnosis at three points along the patient’s ED journey:

  • At triage, o1 gave an exact or very close diagnosis in 67% of cases (when information in the record dump was most limited). The two physicians hit 55% and 50%. 
  • o1 still outperformed the physicians when given all the data collected by the end of the ED encounter.
  • It was only when the physicians were given the most information possible to inform their diagnosis – at the time the patient would have been admitted to the hospital – that the scores finally converged.

The cherry on top? Physician raters couldn’t tell whether the differentials came from o1 or a human. One rater couldn’t tell in 83.6% of cases, the other in 94.4%. 

  • The authors were quick to mention that these results don’t mean AI is ready to replace human physicians. They mean it’s time for rigorous research into how AI can augment care teams, serve as a second opinion, and become a safety layer for clinicians.

The Takeaway

o1 outperforming a couple internists at triage isn’t quite Deep Blue beating Gary Kasparov at chess, but it’s a step in that direction – especially considering OpenAI’s performance jump in just the last week (let alone since o1 launched in 2024).

AI Moves From Proof-of-Concept to Proof-of-Return

Healthcare can cover a lot of ground when it’s moving at the speed of AI, and a new report from McKinsey found that the AI conversation is quickly shifting from proof-of-concept to proof-of-return.

The analysis was based on a survey of U.S. healthcare execs spanning payors, providers, and health services/technology groups.

AI adoption is skyrocketing at all of them. For the first time since McKinsey began tracking the metric in 2023, the orgs that have already implemented GenAI outnumbered those that haven’t.

  • Half of respondents have deployed at least one GenAI use case at their organization, up from just 25% two years ago. Here’s a nice graphic on AI adoption by org type.
  • McKinsey found that leadership teams are no longer questioning whether and where GenAI is relevant, they’re focusing on how it can be used responsibly at scale.

Agents are also building momentum. Despite being the new kid on the AI block, 19% of orgs reported that they’ve deployed agentic AI capabilities.

  • That’s not a huge percentage considering all the new agents we’ve been covering, but another 51% of orgs are actively pursuing agentic AI proofs-of-concept.

Administrative efficiency is the priority. This chart breaks down the areas that respondents see the most potential for GenAI and multiagent systems.

  • 87% ranked administrative efficiency as their leading GenAI use case.
  • 76% said it was also their top priority for multiagent systems.
  • Software infrastructure and engagement trailed as distant contenders for both categories.

Adoption varies by org type. Here’s the overview.

  • Providers are leaning in on clinical productivity (54% are using GenAI to help).
  • Payers are prioritizing administrative efficiency (34%).
  • Health services and tech firms are using GenAI as software infrastructure (52%).

Adoption barriers had more overlap. Across all org types, the chief concerns with GenAI were difficulty integrating with existing workflows, risk/liability, and inaccuracies/bias.

The other shared belief? Nobody implements AI for fun. Everyone expects an ROI.

The Takeaway

AI has arrived in a big way, and McKinsey’s report confirmed that ROI is now the name of the game in every corner of the industry.

Scribes Show Modest Impact at Major Academics

Ambient scribes are back in the spotlight after a new study in JAMA confirmed that they move the needle on productivity metrics, but the jury’s still out on whether that’s the best yardstick for success.

This was a big one. The study examined the impact of AI scribe use on over 1,800 clinicians at five major academic medical centers from 2023 to 2025.

  • The academics: MGB, YNHH, UCSD, UCSF, UC Davis 
  • The scribes: Abridge, Ambience, Microsoft DAX Copilot

Here’s what they found. Clinicians who used AI scribes:

  • Saved 16 minutes of documentation time per eight hours of patient care 
  • Saved 13 minutes of EHR time 
  • Could see one additional patient every two weeks
  • Saw no significant impact on EHR timeoutside of working hours

Usage patterns helped color in the story. While 1,800 AI scribe adopters is one of the largest samples out there, the 6,770 control clinicians were also offered scribes and opted not to use them.

  • The biggest gains went to the biggest users. Clinicians who used the AI scribe for over 50% of visits experienced twice the reduction in EHR time and 3x the reduction in documentation time, yet only 32% of adopters fell into this bucket.

What’s counted? What matters? This isn’t the first study we’ve covered that scores AI scribes based on metrics that researchers can easily measure (EHR time, visits), which isn’t necessarily the same as the metrics that matter most to patients or clinicians.

  • Although this study solidifies that scribes can cut documentation time, the question now is if that time gets reinvested in ways that improve care and outcomes for patients.
  • The results also confirm that the mechanism of action for scribes reducing burnout isn’t through time savings, but it’s still unclear whether it’s from having a couple more moments to take a deep breath throughout the day or from reallocating the extra minutes to things that feel valuable.

The Takeaway

This study offers the most definitive real-world data yet that AI scribes have a modest impact on productivity metrics, but it also confirms that cleaner notes aren’t the only key to improving healthcare experiences.

Qualified Raises $125M to Build AI Infrastructure

In an era of isolated AI pilots, Qualified Health is building the infrastructure to connect the dots.

AI is the star of enterprise transformation. Health systems are looking to deploy and scale AI across their entire organization, and Qualified just raised $125M of Series B funding to make sure every new agent fits into a cohesive constellation.

The core platform has four distinct layers:

  • A data foundation that turns the EHR and external sources into an AI-ready bedrock.
  • A layer that lets hospitals build and deploy AI tools without always starting from scratch.
  • A layer that turns those tools into AI apps and agents deployed directly into workflows.
  • A layer that keeps governance, monitoring, and evaluation at the center of everything.

Qualified doesn’t leave AI to chance. It embeds forward-deployed product leaders alongside health system teams to identify high-priority needs, deploy solutions quickly, and iterate based on actual feedback in the trenches.

That has a couple of major benefits:

  • AI solutions are purpose-built for specific operational problems rather than mass market appeal.  
  • The tight feedback loop allows Qualified to iterate faster than it would be able to with a traditional implementation cycle, which shortens the timescale needed to improve its deployments and demonstrate a measurable impact.

The proof is in the pudding. At the University of Texas Medical Branch, Qualified reportedly generated a $15M measurable run-rate impact within the first six months.

  • That’s an eye-popping number to get on record, and it apparently stemmed from “a real willingness to dive deep” alongside UTMB clinical teams to deploy multiple assistants and automated workflows.
  • Qualified already supports systems representing about 7% of U.S. hospital revenue, and the next chapter is about deepening those partnerships and scaling responsibly.
  • Big ambition also means big competition, and Qualified will be up against everyone from Innovaccer to Epic if it wants to become healthcare’s AI platform of choice.

The Takeaway

Hospitals aren’t looking to AI for incremental improvement. They’re looking to AI to transform how they deliver care, and Qualified just landed another $125M to be the infrastructure that makes that possible.

Google AMIE Shines in First Real-World Study

The gap between benchmark scores and real-world performance has been the theme of the year in AI research, so Google was right on cue with its first prospective clinical trial for AMIE using actual patients. 

Meet the Articulate Medical Intelligence Explorer. AMIE is Google’s flagship “medical AI researcher,” and it teamed up with Beth Israel Deaconess Medical Center to gauge performance in real clinical workflows.

  • 100 patients completed an AMIE interaction before their primary care visit, with AMIE taking medical histories and equipping patients with potential diagnoses to discuss with their PCP.
  • PCPs received the transcript, summary, and AMIE’s management plan prior to the visit. All interactions were monitored live by physicians trained to intervene if safety criteria weren’t met.

AMIE got a gold star. Not only were there zero safety stops across all 100 interactions, patients reported that their attitudes toward AI significantly improved after chatting with AMIE.

  • AMIE’s differential included the correct final diagnosis in 90% of cases (per chart review 8 weeks post-encounter), with 75% top-3 accuracy.
  • PCPs using AMIE reported increased visit preparedness in 75% of cases, as well as potential behavior change in nearly 60%.
  • The quality of AMIE’s differential diagnosis and management plan appropriateness was similar to PCPs, although PCPs won on management plan practicality and cost-effectiveness.

Other findings were less obvious. PCPs had the chart, the physical exam, and the pre-visit transcript, yet AMIE still matched them on differential quality and management safety without taking a single peak at the EHR.

  • That speaks to the ceiling (or lack there-of) for structured AI history-taking, and shows that AI is gearing up to improve patient care in more ways than just making predictions.
  • The fact that PCPs reported better visit preparedness and potential behavior change in over half of cases also highlights how AI can augment – not just replace – clinical reasoning.

The Takeaway

The distance between the bench and bedside is getting shorter, and Google’s AMIE results suggest that conversational AI in primary care is closer to reality than most people might think.

How to Build Patient Trust in Medical AI

AI might move at the speed of trust, but new research in JAMA Network Open shows that trust only moves at the speed of accuracy.

The study had a solid setup. To determine the factors currently driving patient trust in AI, researchers presented 3,000 U.S. adults with a pair of hypothetical AI-assisted visits for a moderate-risk rash. 

  • Each visit had six randomized attributes, such as whether or not a doctor was present, how well the AI performs relative to human clinicians, and various AI governance mechanisms.

AI performance came out on top by a wide margin. Respondents cared more about how well the AI performs than FDA approval, governance, and even having a doctor in the room.

  • The biggest difference came from AI performing better than a specialist, which increased the likelihood of choosing that visit by 32.5%.
  • AI performing at the same level as a specialist boosted visit preference by 24.8%, slightly more than having AI that performs as well as a general practitioner (19.1%).
  • Having an actual doctor present surprisingly only swayed visit preference by 18.4%.

Governance factors also moved the needle. They just didn’t move it much.

  • FDA approval for the AI increased visit preference by a modest 11.1%.
  • Mayo Clinic AI certifications apparently carry just as much weight – also coming in at 11.1%.
  • Local hospital certifications for the AI only gave visits a 7.8% lift.

AI data quality was important. It just wasn’t as convincing as AI performance. 

  • AI that had nationally representative training data boosted visit preference by 11.9%, but it was interesting to see that disclosing bias in the training data had no effect versus not providing any data details.

The written explanations told the same story. Respondents cited AI performance and clinician involvement as the primary reasons for their choices, with many of them expressing comfort with AI as a tool – but not as a standalone decision-maker.

The Takeaway

Widespread AI adoption requires patient trust, and this study did a great job highlighting the specific areas that should be prioritized for building it.

Get the top digital health stories right in your inbox