ARISE’s New MASTerclass on Keeping Up With AI

The folks at ARISE have done it again with their new Medical AI Superintelligence Test (MAST) that enables real-time grading of AI performance across clinical capabilities.

MAST isn’t just a static report. It’s a “living” system that’s built to keep up with AI’s development. The goals are twofold:

  • First, to aggregate the field’s highest quality benchmarks into a shared infrastructure that’s actively maintained.
  • Second, to enable cross-benchmark inference, allowing for more thorough characterization of model performance and the underlying traits which drive it.

The models are evaluated across six domains:

  • Diagnosis, using SCT and NEJM’s CPC.
  • Management, using SCT and NEJM’s CPC.
  • Safety, using NOHARMV2.
  • Radiology, using ReXrank.
  • Multimodal, using MIDAS.
  • Agentic, using MedAgentBench v2 and PhysicianBench.

What would a launch be without a taste of the results? First, MAST breaks down AI agents into clinicians and generalists, allowing fair comparisons with other models in their class:

  • Glass Health’s Glass 5.6 Max took the cake for the clinical models, outperforming all other developers in every category – including diagnostics (79.2%), management (79.5%), and safety (75.8%).
  • Open AI’s GPT 5.5 was the blue-ribbon generalist, nabbing top marks in management (76.3%), agentic capability (56.7%), and safety (73.8%), beating out competitors like Anthropic and Google.  

There’s also a personality test. No, it isn’t Myers-Briggs. This one captures the distinct clinical profiles of the models (small specialist models may outperform larger generalist systems on narrow modalities, while remaining less capable across the broader clinical suite).

  • For example, Google’s Gemini 3.5 Flash was slightly better at radiology and imaging tasks, but didn’t come close to meeting the mark on care management.
  • Meanwhile XAI’s Grok 4 Fast was pretty good at both management and radiology, but gave a middle-of-the-road performance for safety and diagnostic jobs.

So, what’s it mean? Developers no longer have time to study for the next big test when MAST is reevaluating them against every new benchmark. It’ll start being very apparent who’s actually in the lead.

The Takeaway

The ARISE network has done it again. MAST’s live updates match the pace of AI development, giving clinicians real-time feedback on which model to choose for what task.

ARISE Maps the State of Clinical AI

There have probably been hundreds of reports on the medical AI landscape, but there’s only been one State of Clinical AI from the rockstar team at ARISE.

The AI opus delivers the most complete review we’ve seen of a field that’s moving faster than its evaluation practices. It looked at the most influential clinical AI studies from 2025 to answer a trio of important questions:

  • Where does AI meaningfully improve care once it leaves research settings?
  • Where does performance break down?
  • Where do risks remain underexamined?

ARISE brought the heat. The Stanford-Harvard research network produced more highlights than we could count, but here’s a roundup of some of our favorites.

Impressive results in narrow evaluations. AI models have shown “superhuman performance” in research settings, but these results often depend on how narrowly the problem is framed. 

  • In one study, researchers modified standard medical multiple-choice questions so that the correct answer became “none of the other answers.” The clinical reasoning required to solve the question didn’t change. Model performance did. Accuracy dropped sharply across leading AI models, in some cases by over a third.

AI clearly helps prediction at scale. Although diagnostic reasoning was a mixed bag, several studies demonstrated that AI excels at identifying early warning signals from large datasets.

  • A hospital-based study found that a model trained on continuous wearable vital signs predicted patient deterioration up to 24 hours before standard alerts, identifying patients at risk for ICU transfer, cardiac arrest, or death while there was still time to intervene.

Most studies still don’t resemble the reality of healthcare. Clinical work has little to do with answering exam questions, and much to do with reviewing charts, coordinating care, and deciding when not to intervene.

  • A review of 500+ studies found that nearly half of them tested models using medical exam-style questions. Only 5% used real patient data, very few measured whether the models recognized uncertainty, and even fewer examined bias or fairness.

Now what? ARISE offered a few focus areas for 2026 that hit the center of the bullseye for building trust in the latest AI models.  

  • Evaluate models using real-world scenarios to drive evidence-based medicine.
  • Prioritize human-computer interaction design as much as primary outcomes.
  • Measure uncertainty, bias, and harm – especially when it comes to patient-facing AI.

The Takeaway

Healthcare AI has arrived, and ARISE made it clear that innovation won’t be driven by newer models alone. It will depend on whether health systems, researchers, and regulators are willing to apply the same evidence standards to AI that they expect out of any other clinical solution.

Get the top digital health stories right in your inbox