Evaluation of Large Language Model-Generated Recommendations in Glaucoma Surgical Decision-Making


Toprak M., YILMAZ TUĞAN B., YÜKSEL N.

Seminars in Ophthalmology, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1080/08820538.2026.2725223
  • Dergi Adı: Seminars in Ophthalmology
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO)
  • Anahtar Kelimeler: Artificial intelligence, clinical decision support, glaucoma surgery, large language models, ophthalmology
  • Kocaeli Üniversitesi Adresli: Evet

Özet

Objective: To evaluate the utility, rationality, and safety of glaucoma surgery recommendations generated by three prominent large language models (LLMs)–ChatGPT, Microsoft Copilot, and Google Gemini–when applied to real-world clinical scenarios. Methods: Retrospective records from a tertiary hospital were converted into standardized scenarios and stratified into “primary” and “complex” glaucoma groups. Each LLM was prompted to suggest a single surgical approach and provide a rationale. A blinded team of glaucoma specialists evaluated the outputs based on six criteria: appropriateness, rationale quality, specificity, adherence to guidelines, feasibility, and safety risk, using a normalized 0–100 scale. Results: Median overall quality scores across all cases were 80.7 for ChatGPT, 80.7 for Copilot, and 82.7 for Gemini, showing no statistically significant difference in general performance (p =.367). However, case complexity significantly affected performance. For Gemini, appropriateness and rationale quality scores dropped significantly in complex cases and were accompanied by a statistically significant increase in safety risk (p =.009). Although ChatGPT and Copilot demonstrated more stability across groups, their rationale quality was significantly lower in complex scenarios than in primary ones (p =.014 and <0.001, respectively). Pairwise analyses revealed that ChatGPT offered superior rationale quality compared to Copilot, while Gemini exhibited higher specificity. Conclusions: LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.