Nature Medicine, Published online: 10 August 2026; doi:10.1038/s41591-026-04574-5 In piloting and deploying a large language model within a large medical center, we learned that benchmark-based evaluations are insufficient for monitoring and evaluating interactions driven by clinicians, and that this requires new methods for monitoring performance.
Lessons from deploying the ChatEHR system at Stanford Medicine
Michael A. Pfeffer

