Nature Biomedical Engineering, Published online: 24 June 2026; doi:10.1038/s41551-026-01691-x A large-scale benchmark of 87 clinical text tasks across nine languages reveals just how far large language models remain from mastering real-world medical records — and raises the question of what comes next.
From breadth to depth in clinical artificial intelligence evaluation
David Wu

