How to Benchmark Medical AI Agents — PLOS Medicine's Blueprint for Judging Clinical Trajectories
— by Vivax
PLOS Medicine has published "How to benchmark medical AI agents" — a perspective by Sofía Ruhrberg Estévez, Dyke Ferber, Mihaela van der Schaar and Jakob…
PLOS Medicine has published "How to benchmark medical AI agents" — a perspective by Sofía Ruhrberg Estévez, Dyke Ferber, Mihaela van der Schaar and Jakob Nikolas Kather, appearing July 9, 2026. Its argument is compact and overdue: medical AI research is shifting from single-task models toward multimodal, LLM-based agents that operate inside real clinical workflows, and the benchmarks the field relies on were never designed to evaluate that. When an AI system actively gathers information, uses tools, and adapts its decisions as a case evolves, checking only its final output is not evaluation — it is a blind spot. The authors call for benchmarks that assess clinical reasoning, process safety, and resource stewardship along the whole decision path.
Benchmarks do more than measure performance — they define the problem and determine what counts as success. All classical medical AI benchmarks share one evaluation structure: a fixed input, a single response, and a reference answer to score against. But clinical care often allows multiple safe trajectories, and agent competence lives in how information is gathered, how actions are sequenced, and how decisions adapt over time. When the same models are dropped into simulated patient encounters and iterative diagnostic tasks, performance becomes far more variable — a hint that older benchmarks were rewarding static answers, not longitudinal decision-making.
From the paper's framing fall three evaluation dimensions that final-output benchmarks never touch: action appropriateness, process safety, and resource stewardship. The authors propose defining acceptable ranges of practice rather than one fixed answer, with reference trajectories derived from expert consensus, clinical guidelines and simulated patient encounters — designed against reward hacking, with evaluation rigor scaling with agent autonomy. As medical AI moves from prediction to action, evaluation has to move from scoring answers to scoring decisions.