Background This study aimed to evaluate the performance of ChatGPT in identifying hemodialysis (HD) indications from authentic nephrology consultation notes and to compare its recommendations with both real-world clinical decisions and expert nephrologist consensus. Methods This exploratory observational study included 22 anonymized nephrology consultation notes from routine inpatient care at a tertiary care university hospital. Each note was independently evaluated by ChatGPT 5.4 using a standardized zero-shot prompt. The same notes were independently reviewed by three blinded senior academic nephrologists. The majority-vote consensus among these nephrologists was defined as the primary reference standard. Real-world decisions documented by nephrology fellows were evaluated as a secondary comparator. Agreement was assessed using Cohen’s kappa coefficient, and inter-rater reliability among nephrologists was evaluated using Fleiss’ kappa. Results Expert consensus classified eight of 22 cases (40.9%) as requiring hemodialysis and 13 (59.1%) as not requiring HD. Inter-rater agreement among the nephrologists was excellent (Fleiss’ κ = 0.814, p < 0.001). ChatGPT and real-world clinical decisions agreed in 18 of 22 cases (81.8%; κ = 0.633, p = 0.003), whereas ChatGPT and expert consensus agreed in 15 of 22 cases (68.2%; κ = 0.374, p = 0.069). Expert consensus and real-world clinical decisions agreed in 17 of 22 cases (77.3%; κ = 0.553, p = 0.007). Among the nine expert-defined HD cases, ChatGPT agreed in seven cases and demonstrated exact indication-level agreement in four cases. Among the 10 cases classified as HD by both ChatGPT and real-world clinical decisions, exact indication-level agreement was observed in four cases. Discrepancies were most commonly observed in cases involving metabolic acidosis, volume overload, and oliguria/anuria. Conclusion These exploratory findings suggest that ChatGPT may align more closely with nephrology fellows’ decision-making than with senior nephrologists’ judgments. The observed discrepancies may reflect differences in how clinical findings were interpreted and incorporated into the overall clinical context. Although ChatGPT cannot replace expert clinical judgment, it may have potential as an educational tool for trainees, encouraging systematic evaluation of HD indications and structured clinical reasoning. Larger, prospective, multicenter studies are needed to confirm these findings and evaluate ChatGPT’s educational impact in nephrology training.

