A clinical environment simulator for dynamic AI evaluation

Clinical evaluation of large language models (LLMs) currently relies on static datasets and isolated scenarios that fail to capture the cascading effects of healthcare decisions. We propose the Clinical Environment Simulator (CES), a framework that evaluates clinical LLMs within digital hospital env...

Descrizione completa

Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Luyang (Autore) , Kim, Sung Eun (Autore) , Zhang, Xiaoman (Autore) , Kernbach, Julius (Autore) , Kenia, Roshan (Autore) , Acosta, Julian N. (Autore) , Nathanson, Larry A. (Autore) , Haimovich, Adrian D. (Autore) , Rodman, Adam (Autore) , Goh, Ethan (Autore) , Chen, Jonathan H. (Autore) , Shah, Nigam H. (Autore) , Kim, David A. (Autore) , Zou, James (Autore) , Mahmood, Faisal (Autore) , Kather, Jakob Nikolas (Autore) , Lungren, Matthew (Autore) , Natarajan, Vivek (Autore) , Topol, Eric J. (Autore) , Rajpurkar, Pranav (Autore)
Natura: Article (Journal)
Lingua:inglese
Pubblicazione: 12 March 2026
In: Nature medicine
Year: 2026, Volume: 32, Fascicolo: 3, Pages: 820-827
ISSN:1546-170X
DOI:10.1038/s41591-026-04252-6
Accesso online:Verlag, lizenzpflichtig, Volltext: https://doi.org/10.1038/s41591-026-04252-6
Testo
Note sull'autore:Luyang Luo, Sung Eun Kim, Xiaoman Zhang, Julius M. Kernbach, Roshan Kenia, Julian N. Acosta, Larry A. Nathanson, Adrian D. Haimovich, Adam Rodman, Ethan Goh, Jonathan H. Chen, Nigam H. Shah, David A. Kim, James Zou, Faisal Mahmood, Jakob Nikolas Kather, Matthew Lungren, Vivek Natarajan, Eric J. Topol, Pranav Rajpurkar
Descrizione
Riassunto:Clinical evaluation of large language models (LLMs) currently relies on static datasets and isolated scenarios that fail to capture the cascading effects of healthcare decisions. We propose the Clinical Environment Simulator (CES), a framework that evaluates clinical LLMs within digital hospital environments where every decision dynamically alters future states. The CES would use a parallel simulation architecture: a 'hospital engine' that tracks bed availability, staff workloads and equipment status in real time, and a 'patient engine' that simulates disease progression and treatment responses based on LLM interventions. Unlike current benchmarks, the CES framework requires clinical LLMs to execute decisions through realistic electronic health record interfaces, while managing trade-offs between individual patient optimization and system-wide efficiency. The CES enables three critical evaluations absent from current benchmarks: temporal reasoning under evolving constraints, where delayed diagnostics can lead to patient deterioration; resource-aware decision-making, where aggressive workups for one patient may exhaust capacity needed by others; and operational resilience, through adversarial testing with simultaneous emergencies and system failures. By scoring LLM performance on both clinical outcomes and operational metrics, the CES represents a shift toward evaluating clinical LLMs as a dynamic and integrated component of healthcare delivery systems.
Descrizione del documento:Gesehen am 10.06.2026
Descrizione fisica:Online Resource
ISSN:1546-170X
DOI:10.1038/s41591-026-04252-6