A clinical environment simulator for dynamic AI evaluation

Clinical evaluation of large language models (LLMs) currently relies on static datasets and isolated scenarios that fail to capture the cascading effects of healthcare decisions. We propose the Clinical Environment Simulator (CES), a framework that evaluates clinical LLMs within digital hospital env...

Description complète

Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Luyang (Auteur) , Kim, Sung Eun (Auteur) , Zhang, Xiaoman (Auteur) , Kernbach, Julius (Auteur) , Kenia, Roshan (Auteur) , Acosta, Julian N. (Auteur) , Nathanson, Larry A. (Auteur) , Haimovich, Adrian D. (Auteur) , Rodman, Adam (Auteur) , Goh, Ethan (Auteur) , Chen, Jonathan H. (Auteur) , Shah, Nigam H. (Auteur) , Kim, David A. (Auteur) , Zou, James (Auteur) , Mahmood, Faisal (Auteur) , Kather, Jakob Nikolas (Auteur) , Lungren, Matthew (Auteur) , Natarajan, Vivek (Auteur) , Topol, Eric J. (Auteur) , Rajpurkar, Pranav (Auteur)
Format: Article (Journal)
Langue:anglais
Publié: 12 March 2026
In: Nature medicine
Year: 2026, Volume: 32, Numéro: 3, Pages: 820-827
ISSN:1546-170X
DOI:10.1038/s41591-026-04252-6
Accès en ligne:Verlag, lizenzpflichtig, Volltext: https://doi.org/10.1038/s41591-026-04252-6
Accéder au texte intégral
Notes sur l'auteur:Luyang Luo, Sung Eun Kim, Xiaoman Zhang, Julius M. Kernbach, Roshan Kenia, Julian N. Acosta, Larry A. Nathanson, Adrian D. Haimovich, Adam Rodman, Ethan Goh, Jonathan H. Chen, Nigam H. Shah, David A. Kim, James Zou, Faisal Mahmood, Jakob Nikolas Kather, Matthew Lungren, Vivek Natarajan, Eric J. Topol, Pranav Rajpurkar
Description
Résumé:Clinical evaluation of large language models (LLMs) currently relies on static datasets and isolated scenarios that fail to capture the cascading effects of healthcare decisions. We propose the Clinical Environment Simulator (CES), a framework that evaluates clinical LLMs within digital hospital environments where every decision dynamically alters future states. The CES would use a parallel simulation architecture: a 'hospital engine' that tracks bed availability, staff workloads and equipment status in real time, and a 'patient engine' that simulates disease progression and treatment responses based on LLM interventions. Unlike current benchmarks, the CES framework requires clinical LLMs to execute decisions through realistic electronic health record interfaces, while managing trade-offs between individual patient optimization and system-wide efficiency. The CES enables three critical evaluations absent from current benchmarks: temporal reasoning under evolving constraints, where delayed diagnostics can lead to patient deterioration; resource-aware decision-making, where aggressive workups for one patient may exhaust capacity needed by others; and operational resilience, through adversarial testing with simultaneous emergencies and system failures. By scoring LLM performance on both clinical outcomes and operational metrics, the CES represents a shift toward evaluating clinical LLMs as a dynamic and integrated component of healthcare delivery systems.
Description:Gesehen am 10.06.2026
Description matérielle:Online Resource
ISSN:1546-170X
DOI:10.1038/s41591-026-04252-6