Evaluating the accuracy and communication quality of large language models in Ewing sarcoma: a comparative analysis of ChatGPT, Claude, Gemini, DeepSeek, and Grok


Creative Commons License

ÜNYILMAZ C.

Frontiers in Pediatrics, cilt.14, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 14
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3389/fped.2026.1788952
  • Dergi Adı: Frontiers in Pediatrics
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, EMBASE, Directory of Open Access Journals
  • Anahtar Kelimeler: artificial intelligence (AI), Ewing sarcoma, large language models, orthopedic oncology, patient education, pediatric oncology
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Trakya Üniversitesi Adresli: Evet

Özet

Introduction – Large language models (LLMs) are increasingly used to provide medical information, yet their performance in rare pediatric cancers remains largely unexplored. This study aimed to compare the clinical accuracy, comprehensiveness, and communication quality of five widely used LLMs in answering frequently asked questions about Ewing sarcoma. Methods – Twelve representative questions covering diagnosis, treatment, prognosis, and psychosocial support were presented to ChatGPT (GPT-5.2), Claude Sonnet 4.5, Gemini 3, DeepSeek V3.2, and Grok 4. Two orthopedic oncology specialists independently evaluated each response using a 4-point Likert scale assessing clinical accuracy, completeness, clarity, and relevance. Qualitative assessments of empathy and communication quality were also performed. Statistical analyses included the Friedman, Wilcoxon signed-rank, Kruskal–Wallis, and Mann–Whitney U tests. Results – Significant differences were observed among the five LLMs (p < 0.001). ChatGPT achieved the highest overall performance, followed by Claude and DeepSeek. DeepSeek demonstrated the greatest technical accuracy but lower communication quality, whereas ChatGPT provided the best balance between factual correctness and patient-friendly communication. Gemini and Grok produced more superficial responses with lower overall scores. Discussion – Current LLMs can support patient and family education in Ewing sarcoma but should not replace specialist consultation. Although ChatGPT and Claude demonstrated the most reliable overall performance, variability among models remains substantial. Further validation and disease-specific optimization are required before routine implementation in clinical practice.