The folks at ARISE have done it again with their new Medical AI Superintelligence Test (MAST) that enables real-time grading of AI performance across clinical capabilities.
MAST isn’t just a static report. It’s a “living” system that’s built to keep up with AI’s development. The goals are twofold:
- First, to aggregate the field’s highest quality benchmarks into a shared infrastructure that’s actively maintained.
- Second, to enable cross-benchmark inference, allowing for more thorough characterization of model performance and the underlying traits which drive it.
The models are evaluated across six domains:
- Diagnosis, using SCT and NEJM’s CPC.
- Management, using SCT and NEJM’s CPC.
- Safety, using NOHARMV2.
- Radiology, using ReXrank.
- Multimodal, using MIDAS.
- Agentic, using MedAgentBench v2 and PhysicianBench.
What would a launch be without a taste of the results? First, MAST breaks down AI agents into clinicians and generalists, allowing fair comparisons with other models in their class:
- Glass Health’s Glass 5.6 Max took the cake for the clinical models, outperforming all other developers in every category – including diagnostics (79.2%), management (79.5%), and safety (75.8%).
- Open AI’s GPT 5.5 was the blue-ribbon generalist, nabbing top marks in management (76.3%), agentic capability (56.7%), and safety (73.8%), beating out competitors like Anthropic and Google.
There’s also a personality test. No, it isn’t Myers-Briggs. This one captures the distinct clinical profiles of the models (small specialist models may outperform larger generalist systems on narrow modalities, while remaining less capable across the broader clinical suite).
- For example, Google’s Gemini 3.5 Flash was slightly better at radiology and imaging tasks, but didn’t come close to meeting the mark on care management.
- Meanwhile XAI’s Grok 4 Fast was pretty good at both management and radiology, but gave a middle-of-the-road performance for safety and diagnostic jobs.
So, what’s it mean? Developers no longer have time to study for the next big test when MAST is reevaluating them against every new benchmark. It’ll start being very apparent who’s actually in the lead.
The Takeaway
The ARISE network has done it again. MAST’s live updates match the pace of AI development, giving clinicians real-time feedback on which model to choose for what task.

