Ad-verse Effects in Consumer-Facing AI

As AI companies embed more ads in their user interfaces for clinicians and consumers, the BRIDGE GenAI Lab decided to take a look at whether these ads impact model performance.

Turns out, they do. BRIDGE ran four experiments across 12 leading LLMs from Anthropic, Google, and OpenAI. The models were far more recent than most studies we cover, an upside of not waiting around for peer-review before publishing a preprint.

  • Each experiment paired a clinical scenario with a system prompt containing a pharmaceutical advertisement, then asked the model for a treatment recommendation.

Ads definitely moved the needle. Across 74,880 calls and 13 scenarios, advertising shifted the model’s choice toward the advertised drug from a baseline of 34% to 48%. 

  • That’s a jump of +12.7 percentage points on average.

The LLMs had some nice range. Model bias varied widely by developer.

  • Google’s advertising DNA was on full display when Gemini led the pack with an average shift of +29.8 percentage points toward the advertised drug. 
  • Five models from OpenAI were swayed by an average of +10.9 pp.
  • Anthropic’s models were the most resilient at +2.0 pp, and the ever-skeptical Opus 4.6 actually steered away from the promoted drug by -3.8 pp.

Three experiments contrasted three different conditions. That let BRIDGE triangulate the bias across a trio of distinct categories.

  • Equipoise (+12.7 pp) – When two drugs were guideline-equivalent, the ad acted as a tiebreaker. The output was clinically correct, but biased.
  • Suboptimal Drug (+0.6 pp) – When the advertised drug was clinically inferior, models resisted. Only 4.4% of responses chose the suboptimal advertised option.
  • Wellness Supplements (-0.6 pp) – For supplements lacking evidence, endorsement decreased. Anthropic models actively pushed back at -2.4 pp.

The picture was consistent. Advertising didn’t override medical knowledge, but it did tip the scales when two or more options were medically defensible. 

  • Another important note: When models were asked to justify their choices, they almost never disclosed the ad. If they chose the advertised drug, the justification echoed the ad in 52.7% of cases.

The Takeaway

BRIDGE just showed why the real harm with AI advertising might not be patients receiving dangerous drugs. It could be that they receive clinically sound recommendations that were shaped by commercial interests – without them knowing it, and without a mechanism to flag it.

OpenAI o1 Outperforms Physicians on Clinical Reasoning Tasks

A landmark study in Science found that OpenAI’s o1 series outperformed human physicians at multiple clinical reasoning tasks, but that doesn’t mean it’s time to hang up the scrubs just yet.

Researchers at Harvard and Beth Israel Deaconess Medical Center designed the study to evaluate whether LLMs are ready to do what physicians do on a daily basis: review messy patient charts and use that data to determine diagnosis and next steps.

  • They evaluated o1 on clinical cases ranging from patient vignettes to second opinions on 76 real-world ED assessments, which included all the noise and incomplete information that clinicians routinely encounter in the EHR.
  • The refreshingly well-designed study also incorporated a blinded evaluation with two attending physicians at BIDMC and GPT-4.

o1 came to play. On clinical vignettes evaluating management reasoning, o1-preview scored a median of 86%. Not too shabby.

  • It outperformed GPT-4, humans with GPT-4, and humans with conventional resources like UpToDate – all of which scored below 45%.

The ED cases were even more impressive. o1 offered second opinions about the diagnosis at three points along the patient’s ED journey:

  • At triage, o1 gave an exact or very close diagnosis in 67% of cases (when information in the record dump was most limited). The two physicians hit 55% and 50%. 
  • o1 still outperformed the physicians when given all the data collected by the end of the ED encounter.
  • It was only when the physicians were given the most information possible to inform their diagnosis – at the time the patient would have been admitted to the hospital – that the scores finally converged.

The cherry on top? Physician raters couldn’t tell whether the differentials came from o1 or a human. One rater couldn’t tell in 83.6% of cases, the other in 94.4%. 

  • The authors were quick to mention that these results don’t mean AI is ready to replace human physicians. They mean it’s time for rigorous research into how AI can augment care teams, serve as a second opinion, and become a safety layer for clinicians.

The Takeaway

o1 outperforming a couple internists at triage isn’t quite Deep Blue beating Gary Kasparov at chess, but it’s a step in that direction – especially considering OpenAI’s performance jump in just the last week (let alone since o1 launched in 2024).

AI Moves From Proof-of-Concept to Proof-of-Return

Healthcare can cover a lot of ground when it’s moving at the speed of AI, and a new report from McKinsey found that the AI conversation is quickly shifting from proof-of-concept to proof-of-return.

The analysis was based on a survey of U.S. healthcare execs spanning payors, providers, and health services/technology groups.

AI adoption is skyrocketing at all of them. For the first time since McKinsey began tracking the metric in 2023, the orgs that have already implemented GenAI outnumbered those that haven’t.

  • Half of respondents have deployed at least one GenAI use case at their organization, up from just 25% two years ago. Here’s a nice graphic on AI adoption by org type.
  • McKinsey found that leadership teams are no longer questioning whether and where GenAI is relevant, they’re focusing on how it can be used responsibly at scale.

Agents are also building momentum. Despite being the new kid on the AI block, 19% of orgs reported that they’ve deployed agentic AI capabilities.

  • That’s not a huge percentage considering all the new agents we’ve been covering, but another 51% of orgs are actively pursuing agentic AI proofs-of-concept.

Administrative efficiency is the priority. This chart breaks down the areas that respondents see the most potential for GenAI and multiagent systems.

  • 87% ranked administrative efficiency as their leading GenAI use case.
  • 76% said it was also their top priority for multiagent systems.
  • Software infrastructure and engagement trailed as distant contenders for both categories.

Adoption varies by org type. Here’s the overview.

  • Providers are leaning in on clinical productivity (54% are using GenAI to help).
  • Payers are prioritizing administrative efficiency (34%).
  • Health services and tech firms are using GenAI as software infrastructure (52%).

Adoption barriers had more overlap. Across all org types, the chief concerns with GenAI were difficulty integrating with existing workflows, risk/liability, and inaccuracies/bias.

The other shared belief? Nobody implements AI for fun. Everyone expects an ROI.

The Takeaway

AI has arrived in a big way, and McKinsey’s report confirmed that ROI is now the name of the game in every corner of the industry.

Scribes Show Modest Impact at Major Academics

Ambient scribes are back in the spotlight after a new study in JAMA confirmed that they move the needle on productivity metrics, but the jury’s still out on whether that’s the best yardstick for success.

This was a big one. The study examined the impact of AI scribe use on over 1,800 clinicians at five major academic medical centers from 2023 to 2025.

  • The academics: MGB, YNHH, UCSD, UCSF, UC Davis 
  • The scribes: Abridge, Ambience, Microsoft DAX Copilot

Here’s what they found. Clinicians who used AI scribes:

  • Saved 16 minutes of documentation time per eight hours of patient care 
  • Saved 13 minutes of EHR time 
  • Could see one additional patient every two weeks
  • Saw no significant impact on EHR timeoutside of working hours

Usage patterns helped color in the story. While 1,800 AI scribe adopters is one of the largest samples out there, the 6,770 control clinicians were also offered scribes and opted not to use them.

  • The biggest gains went to the biggest users. Clinicians who used the AI scribe for over 50% of visits experienced twice the reduction in EHR time and 3x the reduction in documentation time, yet only 32% of adopters fell into this bucket.

What’s counted? What matters? This isn’t the first study we’ve covered that scores AI scribes based on metrics that researchers can easily measure (EHR time, visits), which isn’t necessarily the same as the metrics that matter most to patients or clinicians.

  • Although this study solidifies that scribes can cut documentation time, the question now is if that time gets reinvested in ways that improve care and outcomes for patients.
  • The results also confirm that the mechanism of action for scribes reducing burnout isn’t through time savings, but it’s still unclear whether it’s from having a couple more moments to take a deep breath throughout the day or from reallocating the extra minutes to things that feel valuable.

The Takeaway

This study offers the most definitive real-world data yet that AI scribes have a modest impact on productivity metrics, but it also confirms that cleaner notes aren’t the only key to improving healthcare experiences.

Qualified Raises $125M to Build AI Infrastructure

In an era of isolated AI pilots, Qualified Health is building the infrastructure to connect the dots.

AI is the star of enterprise transformation. Health systems are looking to deploy and scale AI across their entire organization, and Qualified just raised $125M of Series B funding to make sure every new agent fits into a cohesive constellation.

The core platform has four distinct layers:

  • A data foundation that turns the EHR and external sources into an AI-ready bedrock.
  • A layer that lets hospitals build and deploy AI tools without always starting from scratch.
  • A layer that turns those tools into AI apps and agents deployed directly into workflows.
  • A layer that keeps governance, monitoring, and evaluation at the center of everything.

Qualified doesn’t leave AI to chance. It embeds forward-deployed product leaders alongside health system teams to identify high-priority needs, deploy solutions quickly, and iterate based on actual feedback in the trenches.

That has a couple of major benefits:

  • AI solutions are purpose-built for specific operational problems rather than mass market appeal.  
  • The tight feedback loop allows Qualified to iterate faster than it would be able to with a traditional implementation cycle, which shortens the timescale needed to improve its deployments and demonstrate a measurable impact.

The proof is in the pudding. At the University of Texas Medical Branch, Qualified reportedly generated a $15M measurable run-rate impact within the first six months.

  • That’s an eye-popping number to get on record, and it apparently stemmed from “a real willingness to dive deep” alongside UTMB clinical teams to deploy multiple assistants and automated workflows.
  • Qualified already supports systems representing about 7% of U.S. hospital revenue, and the next chapter is about deepening those partnerships and scaling responsibly.
  • Big ambition also means big competition, and Qualified will be up against everyone from Innovaccer to Epic if it wants to become healthcare’s AI platform of choice.

The Takeaway

Hospitals aren’t looking to AI for incremental improvement. They’re looking to AI to transform how they deliver care, and Qualified just landed another $125M to be the infrastructure that makes that possible.

Google AMIE Shines in First Real-World Study

The gap between benchmark scores and real-world performance has been the theme of the year in AI research, so Google was right on cue with its first prospective clinical trial for AMIE using actual patients. 

Meet the Articulate Medical Intelligence Explorer. AMIE is Google’s flagship “medical AI researcher,” and it teamed up with Beth Israel Deaconess Medical Center to gauge performance in real clinical workflows.

  • 100 patients completed an AMIE interaction before their primary care visit, with AMIE taking medical histories and equipping patients with potential diagnoses to discuss with their PCP.
  • PCPs received the transcript, summary, and AMIE’s management plan prior to the visit. All interactions were monitored live by physicians trained to intervene if safety criteria weren’t met.

AMIE got a gold star. Not only were there zero safety stops across all 100 interactions, patients reported that their attitudes toward AI significantly improved after chatting with AMIE.

  • AMIE’s differential included the correct final diagnosis in 90% of cases (per chart review 8 weeks post-encounter), with 75% top-3 accuracy.
  • PCPs using AMIE reported increased visit preparedness in 75% of cases, as well as potential behavior change in nearly 60%.
  • The quality of AMIE’s differential diagnosis and management plan appropriateness was similar to PCPs, although PCPs won on management plan practicality and cost-effectiveness.

Other findings were less obvious. PCPs had the chart, the physical exam, and the pre-visit transcript, yet AMIE still matched them on differential quality and management safety without taking a single peak at the EHR.

  • That speaks to the ceiling (or lack there-of) for structured AI history-taking, and shows that AI is gearing up to improve patient care in more ways than just making predictions.
  • The fact that PCPs reported better visit preparedness and potential behavior change in over half of cases also highlights how AI can augment – not just replace – clinical reasoning.

The Takeaway

The distance between the bench and bedside is getting shorter, and Google’s AMIE results suggest that conversational AI in primary care is closer to reality than most people might think.

How to Build Patient Trust in Medical AI

AI might move at the speed of trust, but new research in JAMA Network Open shows that trust only moves at the speed of accuracy.

The study had a solid setup. To determine the factors currently driving patient trust in AI, researchers presented 3,000 U.S. adults with a pair of hypothetical AI-assisted visits for a moderate-risk rash. 

  • Each visit had six randomized attributes, such as whether or not a doctor was present, how well the AI performs relative to human clinicians, and various AI governance mechanisms.

AI performance came out on top by a wide margin. Respondents cared more about how well the AI performs than FDA approval, governance, and even having a doctor in the room.

  • The biggest difference came from AI performing better than a specialist, which increased the likelihood of choosing that visit by 32.5%.
  • AI performing at the same level as a specialist boosted visit preference by 24.8%, slightly more than having AI that performs as well as a general practitioner (19.1%).
  • Having an actual doctor present surprisingly only swayed visit preference by 18.4%.

Governance factors also moved the needle. They just didn’t move it much.

  • FDA approval for the AI increased visit preference by a modest 11.1%.
  • Mayo Clinic AI certifications apparently carry just as much weight – also coming in at 11.1%.
  • Local hospital certifications for the AI only gave visits a 7.8% lift.

AI data quality was important. It just wasn’t as convincing as AI performance. 

  • AI that had nationally representative training data boosted visit preference by 11.9%, but it was interesting to see that disclosing bias in the training data had no effect versus not providing any data details.

The written explanations told the same story. Respondents cited AI performance and clinician involvement as the primary reasons for their choices, with many of them expressing comfort with AI as a tool – but not as a standalone decision-maker.

The Takeaway

Widespread AI adoption requires patient trust, and this study did a great job highlighting the specific areas that should be prioritized for building it.

Microsoft Dragon Copilot Gets AI Upgrades

Microsoft might have had the biggest presence at the biggest health IT conference, and it made sure all the lights in Las Vegas were on Dragon Copilot

Unify. Simplify. Scale. Microsoft’s theme at HIMSS was all about making Dragon Copilot a one-stop-shop for information within clinical workflows. It debuted several new capabilities at the show:

  • Integrated medical content from trusted sources
  • Partner-powered AI apps and agents
  • Proactive ICD‑10 specificity suggestions
  • Expanded role-based experiences for physicians, nurses, and radiologists

Partnering is quicker than building. Rather than developing every Dragon Copilot capability in-house, Microsoft has been leaning on outside partners to round out the platform.

  • Dragon Copilot’s clinical evidence feature is a prime example. It brings medical content and other relevant contextual information in-workflow, all curated through new partnerships with Wolters Kluwer, Elsevier, and other vetted sources.

Microsoft Marketplace fills the gaps. It allows users to add AI partner apps directly into their Dragon Copilot workflows. Picture a modular side panel with insights from folks like: 

  • Regard – surfaces comorbidities and relevant diagnoses 
  • Canary Speech – analyzes voice biomarkers for mental health conditions
  • Humata Health – automates prior authorization processes for clinicians 
  • Atropos – generates personalized real-world evidence 
  • Optum – identifies potential coverage issues and supports claims processing 

All roads lead to scribes. When Microsoft first acquired Nuance for $20M back in 2022, it was its second largest acquisition ever behind LinkedIn, and the core offerings were radiology report automation, dictation, and transcription (with humans still pulling a ton of weight).

  • The product formerly known as Dragon Ambient eXperience is now the backbone of Dragon Copilot, and it’s been adding features at a breakneck pace.
  • Microsoft is looking to make Dragon Copilot everything, everywhere, all at once, and so far new partnerships have been the key to making that happen.

The Takeaway

As every digital health company rushes to add scribing to their platform, the OG scribe is rushing to add everything else. Now it just needs to maintain a unified user eXperience.

Infinite Healthcare, What’s It Worth?

Healthcare is one of the few industries where rising usage is treated as a failure, and a16z just published some solid arguments for why that framing might be completely backwards.

Everybody wants to be healthy. The demand for services that help people get and stay healthy is almost limitless, but the supply has always been limited by clinician time and cost.

  • AI balances the equation. It expands our capacity to provide care and drives down its marginal cost, and a16z makes the case that AI opens the door for us to consume an effectively unlimited amount of proactive care – consistent coaching, continuous monitoring, and earlier interventions.

Health is invaluable. As it stands today, when a payor sets reimbursement for a medical service, the rate assumes a certain volume to assess the overall budget for that service.

  • Price x Quantity = Total Medical Expense
  • If AI sends the quantity of the service through the roof while holding the price constant, the total medical expense would skyrocket.

The question isn’t how to avoid this. It’s “what do we get for it?” 

  • Half of all U.S. health expenditures go to 5% of the population, and AI that helps avoid hospitalizations or acute events can generate huge savings from a few patients.
  • Healthier people are also more productive. If AI can help just 1% of the 160M workers in the U.S. work an additional year because they’re healthy, that’s worth $260B in GDP.

How do you price AI for abundant consumption? In a world with truly proactive AI-driven care, delivering more care earlier is what actually bends the cost curve. Pricing shouldn’t punish usage.

a16z looks to other industries as good examples for healthcare:

  • Telecom used to charge for voice and data by the minute because network capacity was scarce, but pricing shifted to unlimited plans as infrastructure improved. Usage went up significantly, but the total market value grew alongside consumption.
  • Music followed the same arc. iTunes sold songs one at a time. Spotify sold access instead. People started listening to more songs, and consumer surplus expanded.

The Takeaway

As AI expands care capacity and access, consumption naturally increases. Affordable access leads to explosions in usage, and business models shift to subscriptions over per-unit pricing. Other industries have made the transition before, and a16z thinks it might be healthcare’s turn.

Amazon Health Connect Sends AI to the Back Office

If the competition for the back office was already hot, it’s a certified wildfire after last week’s debut of Amazon Health Connect

Amazon is pitching Amazon Connect Health as a purpose-built agentic AI solution for the administrative work that gets in the way of care. That’s definitely not fun to read for all the companies that had the same tagline on their booth at ViVE.

It comes with five capabilities straight out of the box: 

  • Patient verification
  • Appointment scheduling 
  • Pre-visit summaries
  • Ambient documentation
  • Medical coding 

What’s the core use case? AWS Director of Healthcare AI Naji Shafi says it’s the entire patient journey.

  • When a patient calls to book an appointment, Amazon Connect Health answers immediately, confirms their identity, checks their coverage, and lines up the visit while they’re still on the line.
  • Before the visit, it reviews their complete medical history across care settings, then surfaces previsit insights like active conditions or trends that might be relevant to closing care gaps.
  • During the visit, it drafts clinical notes for provider review in real-time, with every detail linked back to the moment in the conversation where it was discussed.
  • After the visit, it generates patient-friendly summaries and the medical codes needed for billing, allowing the visit to be payor-ready and submitted within minutes.

But wait, there’s more. Amazon Connect Health integrates natively with Epic, and connects to 100+ EHRs and 35+ HIEs through data integration partners like Redox.

  • It’s also built entirely on AWS HealthLake, the cloud giant’s FHIR data repository that’s now getting new agentic capabilities to help convert records into standard formats.

Early users love it. Amazon One Medical was the perfect sandbox for polishing Amazon Connect Health in clinical settings before opening it to outside partners. It shows in the results.

  • UC San Diego Health is saving a minute per call, diverting 630 hours a week from patient verification to direct support, and slashed call abandonment by 30%.
  • Netsmart’s EHR supports more than 1,300 community provider orgs, and it saw ambient documentation adoption skyrocket 275% – and better staff retention as a result.

The Takeaway

There were already tons of agentic AI solutions competing to automate healthcare’s administrative waste, and now there’s one that’s bankrolled by the biggest bookstore in human history. It’s a crowded space, but $1 trillion per year is also enough bloat to go around.

Get the top digital health stories right in your inbox