Artificial intelligence versus human expertise: reliability of ChatGPT and the London atlas for dental age estimation using panoramic radiographs


PEKER R. B.

BMC ORAL HEALTH, cilt.25, sa.1, 2025 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 25 Sayı: 1
  • Basım Tarihi: 2025
  • Doi Numarası: 10.1186/s12903-025-07360-w
  • Dergi Adı: BMC ORAL HEALTH
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, MEDLINE, Directory of Open Access Journals
  • Trakya Üniversitesi Adresli: Evet

Özet

Background This study evaluated the performance of ChatGPT, a multimodal large language model (LLM), in estimating dental age from panoramic radiographs (PRs) and compared its accuracy and reproducibility with those of the London Atlas (LA) method.
Methods PRs of 620 healthy children aged 6 through 13 years were retrospectively analyzed. An experienced dentomaxillofacial radiologist estimated dental age using the LA, and the ChatGPT-40 model analyzed the same anonymized images to generate automated age predictions. Both methods were repeated after two weeks to assess intra-observer reliability. Predictive accuracy and agreement with chronological age (CA) were evaluated using mean absolute error (MAF), root mean squared error (RMSE), intraclass correlation coefficients (ICC), and Bland-Altman analyses. Statistical significance was set at p<.05.
Results ChatGPT's predictions differed significantly from chronological age (CA), tending to overestimate age in younger children and underestimate age in older children. Compared with the LA, ChatGPT exhibited higher MAE and RMSE values, indicating lower predictive accuracy and greater variability. Error magnitudes were greatest in the 6, 12-, and 13-year-old groups and lowest in the 8-year-old group, whereas the LA showed lower and more stable errors across ages. The LA demonstrated fewer discrepancies and excellent reproducibility (ICC = 0.96) as compared with the moderate agreement of ChatGPT (ICC = 0.703) Overall, the LA provided estimates closer to CA, whereas ChatGPT exhibited greater variability.
Conclusions ChatGPT shows promise for complex decision-making tasks such as dental age estimation, however, its current accuracy, reproducibility, and output stability rernain inferior to established methods such as the LA. The inconsistent predictions observed across repeated evaluations highlight a critical limitation regarding its reliability for clinical and forensic applications. Therefore, ChatGPT-based estimations should be interpreted with caution until future versions achieve more consistent and reproducible performance through population-specific training, model optimization, and multicenter validation.