Background: Multimodal large language models (MLLMs) have emerging potential for interpreting medical images and text, but their performance in orthopedic imaging tasks and the influence of prompt configuration remain insufficiently studied. Objective: This study aimed to evaluate the performance of commercial and open-source MLLMs for diagnosing and staging osteonecrosis of the femoral head (ONFH) and to assess how different prompt configurations affect model performance. Methods: This single-center retrospective diagnostic accuracy study included 159 radiograph patients contributing 318 hip-level observations and 170 magnetic resonance imaging (MRI) patients contributing 340 hip-level observations; 55 patients with both modalities formed the multi-image (MI) subgroup between July 2023 and December 2024. Four MLLMs were evaluated: GPT-4o, Claude 3.7 Sonnet, Qwen2.5-VL 72B, and Gemma 3 27B. Three prompt configurations were tested: single image (SI), image plus radiology description (ID), and MI. Model performance was assessed for ONFH detection; early- versus late-stage differentiation; detailed grading using the Ficat, Association Research Circulation Osseous (ARCO), and Steinberg systems; and grading reliability using intraclass correlation coefficients (ICCs). Results: Model performance varied by prompt configuration and imaging input. For ONFH detection, the SI configuration yielded a mean detection area under the receiver operating characteristic curve (AUC) of 0.55 (SD 0.03), whereas the ID configuration achieved a mean detection AUC of 0.91 (SD 0.01). In radiograph-based ONFH detection, ID input achieved a mean accuracy of 0.88 (SD 0.01); in MRI-based ONFH detection, ID input achieved a mean accuracy of 0.85 (SD 0.01). For early- versus late-stage differentiation, the mean accuracy was 0.65 (SD 0.11) with SI input, 0.78 (SD 0.04) with ID input, and approximately 0.59 (SD 0.10) with MI input. For detailed grading, ID input improved mean accuracy across the Ficat, ARCO, and Steinberg systems compared with SI input. In ARCO grading reliability analysis, the mean MLLM ICC was 0.51 (SD 0.18) for SI, 0.97 (SD 0.02) for ID, and 0.49 (SD 0.15) for MI; surgeon interrater and intrarater ICCs were 0.70 and 0.81, respectively. In the commercial versus open-source model comparison, no significant overall difference was observed between model groups (P=.83). Conclusions: Prompt configuration strongly influenced MLLM performance in ONFH diagnosis and staging. Pairing images with deidentified radiology descriptions improved diagnostic and grading performance, whereas MI input did not provide consistent additional benefit in this retrospective single-center evaluation. These findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.
Prompt Configurations for Multimodal Large Language Models in Diagnosing and Staging Osteonecrosis of the Femoral Head: Multimodel Retrospective Observational Diagnostic Study
Jiesheng Zhu·Pei Fan·Jincheng Shi·Shaoming Chen·Daosen Chen·Zhihan Gao·Yi Lü·T. Yang·Xue Wang·Yimu Lin·Peiyu Xu·Li Li·Xingxing Huang

