*|MC_PREVIEW_TEXT|*

LLM Misinformation, Repeating Prompts, and Software as a Medical Practitioner
February 23, 2026
site logo

Together with

partner logo

“It’s clear that there is indiscriminate investing and most of the AI companies will not succeed. It is also clear that the success story of the near future in every sector including healthcare will be an AI company. Both statements can be true at the same time.”

Suki CEO Punit Singh

Artificial Intelligence

LLMs Still Struggle With Medical Misinformation 

The Lancet Digital Health just published one of the largest-ever stress tests on medical misinformation in LLMs, and it looks like most models still struggle to separate fact from fiction.

Here’s the setup. Researchers probed 20 LLMs with over 3M prompts containing medical information from three different sources: social media posts, simulated clinical vignettes, or real hospital discharge notes with a single fabricated recommendation inserted.

  • Each prompt was presented in multiple versions, once with neutral wording to establish a baseline, then with a series of variations that were emotionally charged or leading.
  • Ten logical fallacies were also used to test how framing influences model behavior, such as appeals to authority (a physician said…) or popularity (everyone agrees that…).

LLMs love fake news. The susceptibility was shockingly high across all models, with the medical misinformation accepted in 32% of the neutral base prompts.

  • That jumped to 46% when the misinformation was embedded in formal discharge notes, but at least the models were more skeptical of the social media content (9%).

Other findings were more counter-intuitive. Eight of the 10 logical fallacies ended up reducing the misinformation acceptance rate rather than increasing it like the authors expected.

  • Only appeals to authority (+2.9 percentage points above the base prompts) and slippery slope prompts (+2.2pp) increased susceptibility, a relatively small impact considering appeals to popularity slashed it by nearly 20pp.
  • Larger models were generally safer, although the language and phrasing had a far greater influence than the parameter count alone. 
  • It was also surprising to see that the medical models performed worse than the general purpose models, with many having weaker lie detectors despite the specialization.

Improving LLM safety is about more than making bigger models. It’s about knowing how information gets presented by actual humans, and having guardrails in place that hold up even when that information is wrong.

The Takeaway

Benchmark performance isn’t real-world performance, and this study provides another reminder that a model’s ability to separate fact from fiction is often more important than its test scores.

Clinician-First Copilot for Value-Based Success

Navina’s AI copilot brings clinical intelligence directly to care teams, turning fragmented data into actionable insights that transform value-based workflows from the back office to the point-of-care. Designed for and loved by physicians, Navina’s Best in KLAS AI reduces missed diagnoses while improving quality metrics and risk adjustment accuracy. Discover how practices are leveraging Navina to enhance VBC performance and improve the clinician experience.

sponsor logo

Abridge & Availity Redefine Payer-Provider Synergy

Abridge is teaming up with Availity to redefine payer-provider synergy at the point of conversation. The collaboration aligns Abridge’s evidence-aware intelligence with Availity’s real-time health information network to create a first-of-its-kind prior authorization experience, with a shared understanding between patients, providers, and payers. Find out how Abridge and Availity are extending conversational intelligence across the revenue cycle.

sponsor logo

The Wire

  • Repeat Your Prompts: A short-but-sweet paper from the Google Research team found that an easy workflow adjustment can improve the output of LLMs: inputting the same prompt twice. The preprint showed that simply repeating the prompt significantly lifted the benchmark performance of leading models (Gemini, GPT, Claude, Deepseek), without increasing the number of generated tokens or latency. This apparently works because the repetition allows LLMs to have the complete context of the prompt before giving its second output, as opposed to building context one word at a time during its first pass.
  • Withings + MedStar: MedStar Health is elevating its Signature concierge medicine service through a new partnership with Withings Health Solutions. Signature patients will now have access to Withings’ BPM Pro 2 connected blood pressure monitor and Body Pro smart scale, giving their clinicians a continuous stream of health data to improve their treatments and personalize care. MedStar designed Signature to offer premium primary care, and it now has premium monitoring devices to make it happen.
  • Physicians Like Epic’s AI Charting: A study in Nature showed that physicians generally find Epic’s AI chart reviews useful, despite the fact that they occasionally omit key details or include blatantly fake ones. UCSD researchers collected feedback from 10 physicians on the quality of 147 AI-generated chart summaries, a new feature that’s currently scaling to EHRs across the country. The feedback was mostly positive even though there were frequent omissions (46), confusing content (20), and hallucinations (5). The authors reached the verdict that AI is ready to augment clinical workflows, as long as clinicians are there to fact check it.
  • It’s Dock Time at Mayo Clinic: Mayo Clinic is rolling out Dock Health’s productivity platform to optimize operational workflows in its cardiovascular, econsult, and specialty contract programs. Dock automates referral processes by streamlining intake, gathering patient history, and improving tracking and outside record requests. That not only gives care teams better visibility from order creation to scheduling, it also accelerates the entire process.
  • Is Not Using AI Unethical? A solid opinion piece in STAT says the answer should be “yes.” The article points to a Nature study that found Google’s AI outperformed six radiologists for mammography screenings, while also reducing false positives and false negatives. Radiologists can be asked to interpret over 100 mammograms in a day, and it’s well-established that marathon shifts aren’t known for improving accuracy. In similar contexts where AI directly targets a known source of preventable harm, the authors say that radiologists not using AI would be like pilots flying without their instruments.
  • Solera Behavioral Health Network: Solera Health debuted its intervention-based Behavioral Health Network to help meet the ever-growing demand for mental health support. The network has some big name launch partners in Calm and Lyra Health, allowing it to encompass both traditional services like talk therapy and lifestyle factors like sleep and stress. The launch also strengthens Solera’s HALO platform, which connects members and employees to benefits based on their unique care needs while allowing payors and employers to manage their point solutions with a single interface. 
  • Software as a Medical Practitioner: An interesting viewpoint in JAMA Internal Medicine makes the case that clinical AI should be regulated more like clinicians, and less like devices. Clinical AI evolves over time, needs ongoing review, and tackles a wide range of tasks. That sounds closer to a clinician than a medical device, and the authors argue that it should be regulated the same way – licensure. That includes initial validation, supervised pilots, a defined scope of practice, discipline boards, and even malpractice accountability.
  • Keragon AI Launch: Keragon announced the launch of Keragon AI to let administrative and clinical operations teams build and deploy HIPAA-compliant automations using plain English. The Keragon platform allows users to integrate hundreds of applications (EHRs, scheduling tools, referral systems, etc.) into their existing workflows or completely automate them, and the new conversational interface removes even more of the technical barriers that have historically slowed healthcare operations.
  • FDA Schedules Decision Support Town Hall: The FDA will hold a town hall for developers of clinical decision support software on March 11 to discuss the agency’s new guidelines for regulating CDS applications. The meeting follows guidance the agency issued in January stating that it would not regulate a CDS application as a medical device if it doesn’t process medical images and HCPs aren’t intended to rely primarily on its recommendations to make decisions.
  • Don’t Touch the Doc’s Comp: A coalition of 38 major healthcare organizations signed a letter backing the Efficiency Adjustment Delay Act (HR 7520). The legislation aims to postpone a 2.5% Medicare reimbursement cut that stemmed from a CMS “efficiency adjustment” until 2030. CMS implemented the reduction based on the belief that technologies like AI simplify procedures, although lawmakers and medical groups argue that the cut ignores clinician burnout and lacks data regarding actual procedure times.

Supporting GLP-1 Weight Loss With RPM

Looking to make GLP-1 prescribing safer and more effective long term? Explore Withings’ suite of remote patient monitoring devices, designed to deliver the continuous, clinically relevant insights care teams need to proactively monitor patients, identify risks early, and intervene with confidence.

sponsor logo

State of Payor Enrollment and Credentialing

AI is changing the way that healthcare leaders tackle provider network management. Medallion’s latest report breaks down the biggest challenges, emerging trends, and how automation is transforming the landscape. Get the insights you need – read the full report today.

sponsor logo

Episode-Based Care: Making TEAM Work

The TEAM model represents one of the largest mandatory reforms in Medicare history, and forward-thinking perioperative leaders are leaning into it. Watch the on-demand recap of C8 Health’s recent fireside chat to explore how episode-based care is reshaping quality improvement – and why the orgs succeeding under TEAM are treating it as a catalyst for transformation, not just a regulatory checkbox.

sponsor logo

The Industry Wire

  1. Trump tariff fallout could affect medical industry. 
  2. White House taps NIH chief to lead beleaguered CDC. 
  3. Anthropic says it’s “critical” to bring company products to EHR
  4. Grail’s cancer blood test falls short in massive U.K. study. 
  5. RCM firms see payor rule changes as biggest threat.
  6. Federal subsidies extend far beyond ACA policies.
  7. United Healthcare raises bar for MA specialist referral.
  8. Hospital operator CHS makes progress on its debt.
  9. Judge strikes down expansion of premerger reporting requirements.
  10. NYC nurse’s strike close to resolution.