AMIE Graduates From Chatbot to Video Visits

The moment has arrived. AI is now as good as physicians at conducting real-time video consultations – so long as patients stick to the script.

There’s a new expert in town. It’s Google’s flagship medical AI researcher named the Articulate Medical Intelligence Explorer, but its friends call it AMIE for short. 

AMIE gets its eyes and ears from a three-agent architecture built on Gemini and Project Astra, which splits the job so no single model has to think and respond at the same time:

  • A Talker agent keeps the conversation flowing with low-latency responses.
  • A Planner agent continuously updates the differential diagnosis and clinical goals.
  • A Perception agent watches the live audio-video stream for clinically relevant cues.

Google put AMIE through the wringer. A randomized study pitted AMIE against 10 board-certified PCPs across 100 telehealth scenarios enacted by 15 professional patient actors, with 20 independent physicians grading the encounters.

AMIE matched or beat the PCPs across the board:

  • Overall clinical rubric score: 83% vs. 68%
  • Top-1 diagnostic accuracy: 91% vs. 77%
  • Perception and examination: 74% vs. 47%

The physical exam gap was the eye-opener. AMIE proactively coached patient actors through maneuvers like range-of-motion tests and self-palpation, while the human PCPs mostly fell back on verbal history-taking.

  • Patient actors preferred AMIE for assessing and explaining their conditions, but still preferred humans for rapport, with interviews describing AMIE as having awkward pauses and odd conversational cadence.

Here’s the disclaimer. These were patient actors following scripts, not real patients. The physicians had to keep their cameras off, and AMIE still whiffed on subtle signals like tremors, nystagmus, and affect – the exact cues where perception matters most.

  • That’s not exactly real-world validation, but that’s already underway through a feasibility study with Beth Israel Deaconess and a nationwide study with Included Health.

The Takeaway

When AMIE first prompted predictions that patients would soon see an “AI doctor” before a human one, the missing physical exam was the go-to rebuttal. That gap just got noticeably smaller, and it doesn’t feel like we’ll be able to write too many more of these before it disappears completely.

AI Will Bend the Cost Curve, for Better or Worse

A new paper in NEJM Catalyst makes the case that AI might finally bend the cost curve in healthcare, just not in the direction patients were hoping for.

“Show me the incentives and I’ll show you the outcomes.” National health expenditures jumped 60% to nearly $5T per year in the last decade, and widely cited estimates from McKinsey suggest AI could shave about $360B off the annual total. 

  • Venrock’s Bob Kocher and Brian Zhao teamed up with USC’s Erin Duffy to explore the areas where those savings will allegedly materialize, and they found the same overarching problem with all of them.

You get what you pay for, and fee-for-service pays for volume. No matter where they looked for potential savings, the authors believe AI is even more likely to have the opposite impact.

  • Drug Development – Faster AI-driven discovery means more new drugs and more eligible patients, but a healthier population doesn’t happen overnight. The net effect for the foreseeable future is more pharma spending, not less.
  • AI Scribes – Under FFS, freeing up physician capacity is a direct path to more visits, which is a hop and a skip away from more fees, tests, referrals, and prescriptions.
  • DTC “AI Doctors” – Could be deflationary if they replace pricier human encounters, or inflationary if every chat ends in an escalation and a testing cascade. Early evidence points to cascades.
  • RPM and CCM – AI dramatically cut the cost of delivering remote care by reducing the clinician time needed to analyze the data, but blanket deployment means more billable monitoring and more incidental interventions (see: United Healthcare’s now-paused move to narrow RPM coverage to two conditions).
  • Admin Automation – The AI-generated savings are real, but in consolidated hospital and insurance markets, that translates to higher margins, not lower prices.

Flip the model, flip the outcome. Every one of those buckets cuts the other way under value-based care arrangements, where the AI force multiplier is more likely to get pointed at complex patients and preventive care rather than volume and more volume.

  • At-risk providers are already using AI to expand access, improve screening, and deliver better treatment. The AI upside is there, it’s just not showing up in the places most people are looking.
  • The authors’ most direct policy fix is simple: more value-based care, and more outcomes-based reimbursement efforts like CMS’s new ACCESS model.

The Takeaway

While AI holds enormous potential to improve healthcare, this paper highlights exactly why it takes more than new tech to bend the cost curve. It also takes the right reimbursement models wrapped around it.

2026 Healthcare Forecast Cloudy, But AI Rays Could Poke Through

Venrock’s 10th annual survey of healthcare insiders reveals they’re a pessimistic bunch lately, harboring cynicism about recent policy developments and the future of health tech IPOs, though views on AI were more of a mixed bag. Let’s break down the results. 

But first, a bit about the survey. More than 200 leaders from all corners of healthcare shared their thoughts with Venrock. Some areas were better represented than others.

  • Respondents skewed toward the private sector (28%), investing (20%), life sciences or pharma (16%), professional services (8%), and academia (7%). 

Venrock loaded up the questionnaire with AI inquiries. Big picture: Insiders are becoming more comfortable with the tech, but remain mindful of its downsides. 

  • Trust in AI grew for 73% and fell for 2% (unclear who hurt them). 
  • Just 9% view AI as the most overrated trend in healthcare.
  • HIPAA breaches (24%), harmful hallucinations (19%), and overspending on healthcare-specific platforms (28%) ranked highest among possible AI pitfalls. 
  • Most expect AI to create an costly arms race between payers and providers (63%) as each side rolls out bots specifically designed to argue with other bots.

Here’s another fun one: M&A targets. There’s no consensus on who will get snapped up next, but the industry seems confident it will be a big name in AI-powered services. 

  • OpenEvidence (21%), Komodo Health (18%), Abridge (14%), and Sword Health (11%) were the top answers out of 10 companies, but only after none of the above (46%).

So the hottest firms are going public? Nope — that’s one thing people can agree on. 

  • Only 3% think health tech IPOs will be back in style this year, with the rest split roughly down the middle between 2027 and 2028 or beyond. 
  • For those keeping score, 42% of last year’s respondents predicted a health tech firm would go public in the first half of 2026. Tumbleweeds…

In fairness, it’s hard to predict the future, especially with $1.15T in Medicaid cuts looming over everyone’s heads. 

  • Will they harm rural hospitals? Empty state coffers? Ruin MCOs? Strain blue-state safety net hospitals? Most checked all of the above (65%). 

The Takeaway

Some of the smartest folks in healthcare think we’re heading toward a world of payer-provider bot wars, sluggish health tech IPOs, and brutal fallout from Medicaid cuts. Here’s to hoping Venrock’s survey missed the mark. 

IntelePeer Sister Company Aqurio Enters the AI Agent Fray

Telecommunications company IntelePeer thinks it’s in the right place at the right time to capitalize on rising interest in agentic healthcare AI with the launch of sister company Aqurio. 

IntelePeer cut its teeth during the voice over internet protocol boom of the 2000s, eventually moving toward cloud-based communications services and now AI. 

  • As AI took up more of its focus, IntelePeer decided it was time for a new organization, leading to the creation of Aqurio. 

So what does IntelePeer know about healthcare? Quite a bit, since many of its automated customer service tools are used by providers and health plans. 

  • Through this experience, IntelePeer became keenly aware of the industry’s operational pain points, like billing backlogs and call center wait times. 

But why launch Aqurio now? IntelePeer was convinced by recent advancements in AI. 

  • Specifically, AI can now tackle end-to-end tasks in regulated environments, many of which once took large teams and lots of resources to coordinate. 
  • AI has also become cheaper, with inference costs falling over 95% since 2022. 
  • Wait any longer, and another firm might solve the problems Aqurio is after. 

Aqurio is rolling with three main products, but one platform. By keeping agents unified, the company aims to stop patients and revenue from falling through the cracks. 

  • For administrative tasks, SmartAgent answers calls, texts, and chat messages to support things like scheduling, insurance verification, and billing inquiries.
  • When it comes to outreach, SmartEngage sends collections and patient recall messages on the provider’s behalf. 
  • On the data front, SmartAnalytics combs customer service interactions for insight into KPIs, ROI, human agent performance, and behavior of the other two agents. 
  • As for clinical follow-up, SmartCare handles visit summaries, post-surgical assessment, and symptom flagging.

The agentic AI market may be crowded, with no shortage of well-funded companies, but Aqurio could hit the ground running by leveraging IntelePeer’s foundation. 

  • IntelePeer has logged more than 1B customer interactions. 
  • Aqurio’s platform has HIPAA, HITRUST, SOC 2 Type II, and other certifications out of the gate. 

The Takeaway

With the creation of Aqurio, IntelePeer is aggressively pushing into agentic healthcare AI without sacrificing its core telecommunications business. Time will tell if Aqurio can muscle out the competition, but IntelePeer’s more than 20-year history gives its sister company a massive leg up over three-guys-in-a-garage startups.

New Model Predicts 900 Diseases From Real Records

Last week brought a potentially significant step forward for early diagnosis in the form of a new AI model that can predict a patient’s next diagnosis across nearly 900 diseases using just real-world medical records.

Meet DT-Transformer. Researchers at Harvard, Brigham and Women’s, and the Broad Institute unveiled the GPT-style foundation model in a new arXiv preprint.

  • DT-Transformer reads a patient’s medical history as a sequence and predicts which disease will show up next, and when it might come knocking.

The real headline is the training set. DT-Transformer was trained on 57.1M structured EHR entries from 1.7M patients across MGB’s 11 hospitals and 200 clinics (2000 to 2024) – the messy reality of U.S. clinical care rather than a polished research data set.

The numbers were impressive. A few standouts:

  • DT-Transformer achieved a median AUC of 0.871 across 896 disease categories, with AUC over 0.5 for every condition.
  • The model crushed an age- and sex-based baseline by +0.214 AUC (0.871 vs. 0.657), beating it on 96% of diseases.
  • All of that was accomplished with a featherweight 2.2M parameter model that’s small enough to run just about anywhere (by comparison, Claude Fable 5 has about 6 trillion parameters).

The real test told a humbler story. When the team ran a true prospective test forecasting new diagnoses DT-Transformer had never seen, median AUC slipped to 0.713.

  • That still beats the baseline on 80% of diseases, but that gap between the retrospective flex and the prospective reality is fairly significant.  
  • A 0.871 headline number and a 0.713 crystal ball aren’t the same product, and the second one is the one that patients would actually have to deal with.

One other highlight worth mentioning: including every repeated diagnosis worsened model performance, “drowning out the signal” rather than improving predictions – another reminder that more data isn’t always better. Better data is.

The Takeaway

Population-scale risk forecasting that runs on a model smaller than most phone apps is a real milestone, and training on routine records instead of a spotless biobank is exactly the kind of thing that could get models like DT-Transformer in front of actual patients – assuming the data holds up to peer-review.

UpDoc Lands First FDA Clearance for Patient-Facing AI 

UpDoc just landed the first FDA clearance for a patient-facing AI model, which acts as a “concierge doctor” to support patients between visits. Good news for the human docs reading this – it isn’t going after your job just yet.

What’s UpDoc? It’s a clinical AI platform that unifies clinical guidelines, longitudinal patient context, and physician governance to safely execute real-world care workflows.

What isn’t UpDoc? An AI doctor.

  • The 510(k) clearance had a narrow scope. It allows the AI to call or message patients between visits and adjust their insulin doses within parameters set by human clinicians.
  • UpDoc says its AI will ease doctors’ workloads and help patients better manage illnesses like Type 2 diabetes. 

The data backs that up. A study in JAMA Network Open saw 32 patients with T2D randomized to receive support from UpDoc (daily voice AI check-ins to record blood glucose and adjust insulin) or standard care (AKA log their own data until they see their doctor in person).

  • The AI group hit their target blood glucose in 15 days, compared to the standard care group where less than half got there at all within the 8 week study period.
  • That trial provided the clinical foundation for the now-cleared solution, which is set to be piloted at Cleveland Clinic, UCSF Health, and Allegheny Health Network.

UpDoc is taking the road less traveled. It’s not the only AI startup in this wheelhouse, but so far it’s one of the only ones that doesn’t seem to be actively avoiding FDA regulation. 

  • The most notable example is Doctronic, which has been testing its AI prescription tech through a state-run program in Utah rather than seeking a full-fledged authorization.
  • That’s an easier path to market than vaulting over FDA hurdles, but it doesn’t get you the “world first” feather in your cap that now belongs to UpDoc.

So, now what? The FDA has long debated how to regulate AI, and UpDoc could be the first sign that they’re getting comfortable enough with the tech to give the green light to more models.

  • With the first clearance out of the way, other AI developers also have an established precedent and a blueprint to follow suit.

Don’t forget about the docs. Besides the regulatory shakeout, it’ll be equally interesting to see how this new breed of AI ends up in the hands of clinicians.

  • We saw OpenEvidence fold a new biomarker for heart disease into its platform just last week, and the most direct path to a wide distribution for many soon-to-be-cleared AI tools could be similar licensing partnerships.
  • Plenty of companies already have a massive user base and are actively expanding the clinical scope of their platforms – Abridge, OE, Doximity, the list goes on for a while. It feels like licensing models from the UpDocs of the world is a natural next step after all the journal partnerships we’ve been seeing now that FDA clearance is part of the picture.

The Takeaway

The FDA finally cleared its first patient-facing clinical AI model, and UpDoc might have been the first domino, but it definitely won’t be the last.

Assort Closes $120M to Scale Voice AI Across Healthcare

If you needed any more proof that communication friction is one of the biggest pain points for patients and providers, look no further than Assort Health’s just-closed $120M Series C – its third funding round in 18 months.

Assort started with a simple thesis. Unlock the front door of healthcare, and the rest will follow. Assort originally aimed its voice AI agents at scheduling because it meant solving for two key ingredients needed to solve everything else downstream: 

  • The care protocols required to handle that first interaction.
  • The patient communication data that flows into the rest of the journey.

The first call is an important moment. Mistakes here mean the patient never comes back, and Assort’s edge in preventing that is its Synapse agentic model.

  • Synapse learns specialty workflows across every deployment, then simulates the edge cases to stress test them before any agents go live.
  • That allows even non-technical teams to safely implement Assort’s agents at scale, which fuels an AI development flywheel that’s already learning from 190M patient interactions, 62M care protocols, and 1.6M decision pathways.

Assort covers the entire patient journey. What began as the first voice AI agent to schedule a specialty appointment has grown into a full-fledged voice AI platform that includes:

  • Concierge – handles inbound calls, triage, lab requests, med refills, scheduling, eligibility checks, and intake.
  • Activate – reaches patients proactively to close referral loops and act on care gaps, recover no-shows, and resolve payments.
  • Orchestrate – runs the operational work behind each visit and writes every detail back to the EHR.
  • Empower – equips staff with an AI copilot to manage complex patient access needs in real time.

Patient Journey Memory ties it all together. The capability is built on three pillars.

  • Each patient gets a personal AI agent that knows their context and preferences so they don’t have to keep repeating the same story every time they interact with their provider.
  • The agents share the same data and talk to each other, so care gaps surface wherever the patient happens to engage.
  • Having a continuous journey across every interactions allows the platform to activate patients when they’re high intent.

Next stop: everywhere. Every new tech generation sees a flood of new solutions, then only a few survive. Voice AI is about to hit that same shakeout, and Assort plans on sticking around.

  • The funding was earmarked for bringing on veteran C-levels to make that happen, and expanding into health systems ranging from community-based organizations to the biggest academic medical centers in the country.
  • Major systems like John Muir Health are already signed on as demand grows for platforms that can support increasingly complex ambulatory operations – the exact kind Assort is uniquely tuned to solve.

The Takeaway

Assort is looking to become the voice AI transformation partner for every healthcare provider in the country, and if its funding tempo is any indication, it’s moving with enough urgency to actually pull it off.

New Studies Show AI Outperforms Physicians, Just Not at Medicine

In case last week’s AI drama wasn’t hot enough, a pair of new studies in Nature cranked up the heat by finding that AI agents beat physicians on ER and care management tasks – just not real ones.

“Towards autonomous medical artificial intelligence agents.” The first study took a look at MIRA, an AI agent developed in Germany that operates inside a sandboxed EHR environment.

  • Using 574 real emergency department cases, researchers had MIRA chat with another patient agent and execute entire care workflows, such as investigating diagnoses, ordering labs, and triaging for hospital admission. 

The headline: MIRA significantly outperformed four board-certified physicians. The agent had higher overall diagnostic accuracy (87.8% vs. 78.1%), was better at ordering correct procedures like laparoscopic appendectomy (53.5% vs 38.3%), and had 35% better guideline alignment.

The reality: ER doc Graham Walker, MD, put it perfectly on LinkedIn: “There is no way in hell that humans mismanaged almost 30% of appendicitis cases, the most common ‘surgical emergency’ that we’ve all seen hundreds of in our career.”

  • It turns out the EHR sandbox needed 21 keystrokes to get this right, and the physicians failed unless they explicitly searched and entered a “laparoscopic appendectomy.” AI is built for that, humans not so much.

“Towards conversational AI for disease management.” The second study explored whether Google’s AMIE agent could expand from pure diagnostics to longitudinal care management.

  • The blinded study pitted AMIE against 21 primary care physicians on 100 multi-visit cases, with the agent pulling live guidelines and drug references to produce structured management plans.

The headline: AMIE’s care plans were better than PCPs across the board. The agent notched higher marks on management reasoning, precision of investigations, and guideline alignment.

The reality: AMIE operated in a world without prior auths, without formulary restrictions, and without social needs that patients didn’t want to bring up. The authors didn’t pretend otherwise.

The Takeaway

This might sound familiar, but these studies show that MIRA and AMIE performed well in ideal scenarios, not in the messy trenches of real-world medicine. That said, the results aren’t important because AI beat a benchmark, they’re important because AI took another big step toward “delivering actions” instead of just “delivering answers.”

General-Purpose LLMs Outperform Healthcare-Specific Models

We might have just gotten our spiciest study of the year after new findings in Nature Medicine showed that general-purpose LLMs outperform specialized healthcare models straight out of the box.

It was a battle of the bots. Researchers pitted OpenEvidence and UpToDate Expert AI against three frontier models that anyone with a web browser can pull up in two seconds: GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6.

The models were tested across three domains:

  • medical knowledge (MedQA)
  • expert clinician alignment (HealthBench)
  • 100 real physician queries (RCQ) scored by 12 blinded clinicians  

It was a clean sweep. The general-purpose LLMs outperformed the specialized models on all three evals, and by a healthy margin. This chart gets the point across.

  • On MedQA, Gemini led the pack with 97.4% accuracy (vs. 89.6% for OE and 88.4% for UTD). Fun fact, the frontier models had a huge advantage here since their training data included these exact questions (and answers).
  • On HealthBench, GPT-5.2 dominated with an 88%. It’s almost like OpenAI invented the benchmark.
  • The RCQs were probably the most clinically meaningful component, and all three frontier models took the podium here as well. It was a bit odd that the researchers didn’t share the specific questions, and OE definitely thought so too.

OpenEvidence hit back hard and fast. It went straight to its socials to let the world know that the study was not only poorly designed and biased, but that the authors had reached out for API access to help build a competing product. Request denied.

  • OE also pointing out the training data contamination issue with MedQA, and critiqued HealthBench for scoring responses based on subjective stylistic choices (in one example OE scored 20% “worse” because it didn’t use a specific email header).
  • The cherry on top was OE revealing that the real-world clinician queries were only added after peer reviewers flagged the study for having weak evidence. Big if true.

Obligatory disclaimer: the models were evaluated back in February, and the performance gap could easily be even wider today. 

The Takeaway

OpenEvidence and UpToDate didn’t become successful by being better AI developers than OpenAI and Anthropic. They did it by doing the things that don’t show up in benchmarks – curating sources of verifiable evidence, wrapping them in an interface that docs actually enjoy using, and earning their trust one question at a time. If anything, this study confirmed that those matter now more than ever.

Abridge Unveils New Platform, Teams Up With Lilly and Nvidia

Patients, platforms, Lilly, and Nvidia. Abridge’s first keynote had it all.

There were enough major announcements to fill an entire issue of DHW, so here’s the abridged version of the top stories to come out of NYC.

The new platform stole the show. Abridge unveiled “the first AI-native clinician intelligence platform” organized around patients, built for clinicians, and designed to help health systems.

  • Before the visit: The platform surfaces care gaps and relevant clinical context so clinicians can address what matters during the visit instead of discovering it in retrospective chart reviews.
  • During the visit: Abridge suggests discussion topics while delivering evidence-based answers to clinical questions from a growing content library that includes new specialty-focused partners like AAFP, AAN, ADA, and ASCO.
  • After the visit: Abridge generates documentation, flowsheets, patient summaries, orders, and billing codes (soon to be fine-tuned through a new partnership with AHIMA).

“The base unit of healthcare is a clinician caring for a patient.” As Abridge pushes into new models of care delivery, its platform will provide the connective tissue between the clinical workflows where care actually gets delivered and outside orgs like payers or life sciences firms.

  • The keynote highlighted some key examples: Cigna was on stage discussing how embedding AI in clinical workflows has the potential to unlock real-time claims adjudication, and Aetna shared how it could help realize the promise of VBC.
  • More than 300 health systems are already live, including a just-announced rollout at Northwestern Medicine.

Eli Lilly is buying into the vision. The pharma giant made a strategic investment in Abridge’s next chapter, and even though the keynote was light on details, the move started to add up after seeing one of the new capabilities coming to the platform: clinical trial screening.

  • By comparing clinical guidance with patient-provider conversations in real-time, Abridge can surface relevant trials directly in the encounter – the moment it matters most. 
  • They didn’t mention a check size, but big opportunities attract big investments, and identifying candidates while initiating screening at the point of care sounds huge.

Last, but certainly not least, Nvidia. Abridge is teaming up with Nvidia to develop a first-of-its-kind foundation model for clinical conversations that’s trained, shaped, and evaluated against real-world conditions.

  • We’ll have to wait until later this year to see it in action, but a little pre-, mid-, and post-training magic with Abridge’s de-identified clinical data will apparently help make it the first model that can “reason clinically from its foundation.”

The Takeaway

If the keynote made one thing crystal clear, it’s that Abridge’s platform doesn’t revolve around AI documentation. It revolves around patients, and every new feature is purpose-built to prove it.

Get the top digital health stories right in your inbox