*|MC_PREVIEW_TEXT|*

Clinical AI vs the World, OpenLoop Launchpad, and Implementation Struggles
By Jason Barry
June 18, 2026
site logo

Together with

partner logo

“The benchmark crowns a model. It doesn’t sign the chart.”

Mankato Clinic CMO Andrew Lundquist

It was tough to decide whether to tune into the World Cup or the Clinical AI soap opera this week. Get your popcorn ready.

Jason

Artificial Intelligence

General-Purpose LLMs Outperform Healthcare-Specific Models

We might have just gotten our spiciest study of the year after new findings in Nature Medicine showed that general-purpose LLMs outperform specialized healthcare models straight out of the box.

It was a battle of the bots. Researchers pitted OpenEvidence and UpToDate Expert AI against three frontier models that anyone with a web browser can pull up in two seconds: GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6.

The models were tested across three domains:

  • medical knowledge (MedQA)
  • expert clinician alignment (HealthBench)
  • 100 real physician queries (RCQ) scored by 12 blinded clinicians  

It was a clean sweep. The general-purpose LLMs outperformed the specialized models on all three evals, and by a healthy margin. This chart gets the point across.

  • On MedQA, Gemini led the pack with 97.4% accuracy (vs. 89.6% for OE and 88.4% for UTD). Fun fact, the frontier models had a huge advantage here since their training data included these exact questions (and answers).
  • On HealthBench, GPT-5.2 dominated with an 88%. It’s almost like OpenAI invented the benchmark.
  • The RCQs were probably the most clinically meaningful component, and all three frontier models took the podium here as well. It was a bit odd that the researchers didn’t share the specific questions, and OE definitely thought so too.

OpenEvidence hit back hard and fast. It went straight to its socials to let the world know that the study was not only poorly designed and biased, but that the authors had reached out for API access to help build a competing product. Request denied.

  • OE also pointing out the training data contamination issue with MedQA, and critiqued HealthBench for scoring responses based on subjective stylistic choices (in one example OE scored 20% “worse” because it didn’t use a specific email header).
  • The cherry on top was OE revealing that the real-world clinician queries were only added after peer reviewers flagged the study for having weak evidence. Big if true.

Obligatory disclaimer: the models were evaluated back in February, and the performance gap could easily be even wider today. 

The Takeaway

OpenEvidence and UpToDate didn’t become successful by being better AI developers than OpenAI and Anthropic. They did it by doing the things that don’t show up in benchmarks – curating sources of verifiable evidence, wrapping them in an interface that docs actually enjoy using, and earning their trust one question at a time. If anything, this study confirmed that those matter now more than ever.

Any Use Case, Any Specialty

Bunkerhill’s Carebricks platform doesn’t stop at surfacing insights. It translates them into real-world action. From automating prior auths to closing care gaps, Carebricks lets health systems design and deploy AI agents for any clinical or operational need – without adding to anyone’s manual workload. Learn how Carebricks can automate actions for your patients today.

sponsor logo

Redefining Patient Monitoring for Obesity & Diabetes

Weight management programs live and die by adherence. Discover why partners like Calibrate,
9am Health, and Wondr Health trust Withings to keep members weighing in, with
cellular-connected scales that deliver instant weight insights, full body composition data, and the
lowest possible barrier to action throughout their weight loss journey.

sponsor logo

The Wire

  • OpenLoop Launches Launchpad: OpenLoop kept its hot streak alive with the debut of Launchpad, a turnkey platform that lets organizations launch a virtual care brand in as little as 24 hours. Launchpad brings together patient acquisition, intake, care delivery, provider operations, and analytics into one complete platform, which is about as close to telehealth-in-a-box as anyone’s ever gotten. CEO Jon Lensing shared some great insight into the mission behind the momentum on a recent episode of the DHW Show.
  • Everyone’s Bullish on AI: Arcadia put out a survey that suggests healthcare leaders are more bullish than ever on AI’s potential, but they’re having a hard time translating promises into results. The poll found that 52% of healthcare leaders believe AI can transform healthcare when applied correctly, yet only 14% report AI insights are fully integrated into day-to-day decision-making. That said, it seems like the debate over whether AI belongs in healthcare is over, with just 6% of execs viewing the tech as overhyped and introducing more risk than value. 
  • Tom Symptom-Checker: Lumeris took the lid off a new symptom-checking capability within its Tom AI Primary Care-as-a-Service platform. The feature is powered by Google Gemini Flash to let high-risk patients and caregivers report symptoms and ask health questions between scheduled visits, helping care teams identify and respond to emerging needs before they escalate. The capability is already in limited deployment with a leading Medicare Advantage plan, and it sounds like broader availability is on the way later this year.
  • AI Authorizations: The FDA updated its list of authorized AI-enabled medical devices through the end of Q1 2026 and provided a decent look at where the AI momentum is heading. The list now includes data through the end of March 2026, and shows that the FDA has authorized 1,524 AI-enabled medical devices since it began keeping track in 1995, up 5.1% QoQ. In the first quarter of 2026, the FDA authorized 92 new devices, 28% more than Q4 of last year. The Imaging Wire has you covered with the full breakdown. 
  • Future Health Index: Philips’ always-excellent Future Health Index showed that over half of clinicians are now using AI to reclaim an average of 132 hours per year. The massive survey of 2k healthcare professionals and 20k patients confirmed that AI adoption is outpacing readiness, with 64% of clinicians turning to personal AI tools when workplace options don’t meet their needs. The majority of clinicians are also starting to view AI as a bonafide teammate, with 65% agreeing that AI agents will support clinical reasoning and 82% expecting their role to be more focused on higher-value clinical work.
  • Coding Quality Council: CodaMetrix joined forces with five health systems to form a first of its kind governance body for developing quality standards around AI medical coding. Statistics like “95% accuracy” are commonly cited as gold standard results in AI-powered autonomous medical coding with no consistent, industry-wide definition or method for measuring that accuracy, even when coder agreement rates hover around 50% even among certified coders. The goal of the new Coding Quality Council is to build an objective framework for measuring coding quality so patients pay for the care they receive – not more, not less.
  • InStride Series C: InStride Health locked in $30M of Series C funding to give young patients better access to behavioral healthcare. The insurance-covered specialty care startup tackles complex anxiety, OCD, and related disorders across 17 states, and the fresh funding was earmarked for more expansion. InStride’s model combines evidence-based specialty care with purpose-built AI tools that help clinicians deliver consistent care while keeping decision-making firmly in human hands.
  • Compounding Crackdown: The FDA is cracking down on compounded GLP-1s, and 25 telehealth companies recently received strongly worded letters over their claims about the weight-loss drugs. The letters were sent to companies like Medica Weight Loss, Ready Med, Clover Meds among others earlier this month, warning them about false or misleading claims (mainly positioning their products as the same as ​approved GLP-1s). The agency proposed ​excluding Novo Nordisk and Eli Lilly’s weight-loss drugs from a key compounding list in April, which would potentially limit large-scale ​production by outsourcing facilities.

How MUSC Is Bringing Care Closer to Home

MUSC Health teamed up with Ovatient to accomplish a simple goal: ensure access to care isn’t determined by a ZIP code. A third of South Carolina residents live in rural areas, but travel barriers don’t make preventative care any less crucial. See how Ovatient’s virtual-first approach is improving access to urgent care, primary care and integrated behavioral health and reducing leakage – without MUSC providers doing the heavy lifting.

sponsor logo

Abridge Named #1 Best in KLAS – Again

KLAS just named Abridge #1 Best in KLAS for Ambient AI for the second year in a row. The recognition was based on direct customer feedback from the nation’s largest and most complex health systems, which gave Abridge the highest overall satisfaction score and A+ ratings across Culture, Loyalty, Relationship, and Value. Discover why Abridge is the market-leading AI platform for clinical conversations.

sponsor logo

Privia Accelerates VBC Success With Navina

How do you give physicians new AI tools to accelerate VBC without slowing them down? Privia Health found its answer with Navina. They co-designed an under-one-hour training program that onboarded 800+ clinicians in the first year, driving 87% weekly active usage while providing clearer visibility into their patient panels. Read the full case study to see what it takes to make adoption stick and outcomes follow.

sponsor logo

The Resource Wire

  • Evidence, in the Flow of Care: Heidi brings trusted guidelines and peer-reviewed research directly into clinical workflows so decisions don’t stall care. Clinicians get clear, evidence-based answers without leaving the conversation. No ads, no limits, and no outside interests getting in the way of care. Find out how with Heidi Evidence.
  • Same Tool, New Name, Better AI: DoxGPT is now called Ask, and it’s powered by a new agentic reasoning engine that delivers better, faster, and more reliable responses. Physicians can still find verified answers to complex clinical questions, integrated drug references, and full-text access to over 2,000 top journals – all in the Doximity workflows they’re already using every day. Don’t wait, Ask.
  • State of Payor Enrollment and Credentialing: Over half of provider orgs are losing revenue due to credentialing delays – with many missing out on over $1M annually. Medallion’s new report unpacks the forces quietly undermining operational and financial performance, and how leaders across the industry are addressing them. Head over to the full report to get insights tailored to your role and org type.
  • AI That Cares: Healthcare isn’t one-size-fits-all. Your AI shouldn’t be either. Tucuvi works with health systems to deliver personalized care to their unique populations, from routine administrative tasks to the most complex clinical workflows. See why leading healthcare organizations trust Tucuvi to support their care teams. 

The Industry Wire

  1. NIH moves to reduce animal testing in research.
  2. SpaceX IPO reveals healthcare ambitions.
  3. Epic forces staffing firm to rebrand in trademark settlement.
  4. Medicare trust fund to run out by 2033.
  5. OhioHealth settles antitrust suit with the DOJ.
  6. HHS ramps up information blocking enforcement.
  7. Anthropic wants govt to shut down AI that threatens hospitals.
  8. FBI stages fake hospital to train for ransomware attacks.
  9. TN-based Lifepoint Health taps new COO.
  10. FDA clears first over-the-counter CGM for children.