Evaluating the accuracy of ChatGPT and Gemini models in dental specialty exams: A comparative study of pediatric and restorative dentistry

Ayşenur Altuğ Yıldırım(1), Menekşe Alim(2)
(1) Çankırı Karatekin University, Faculty of Dentistry, Department of Restorative Dentistry, Çankırı, Türkiye,
(2) Private practice, Ankara, Türkiye

Abstract

Aim: The integration of large language models (LLMs) such as ChatGPT and Gemini into healthcare education presents novel opportunities for knowledge dissemination and assessment. Despite their growing prominence, empirical evidence regarding their reliability and domain-specific accuracy in dental specialty examinations remains limited. This study aimed to evaluate and compare the performance of four AI-driven LLMs—ChatGPT-3.5, ChatGPT-4o, Gemini 2.0 Flash, and Gemini 2.0 Advanced— by employing official past questions from the Turkish Dental Specialty Examination (DUS) in Pediatric Dentistry and Restorative Dentistry.


Methodology: A total of 178 multiple-choice questions (90 in Restorative Dentistry and 88 in Pediatric Dentistry) were obtained from DUS examinations conducted between 2012 and 2017. All items were presented in Turkish to each model via standardized prompts in isolated sessions to prevent cross-interaction and memory bias. Response accuracy was determined by evaluating the correctness of the chosen answers. Comparative analyses between models, domains, and annual trends were conducted using non-parametric statistical tests, including the Mann–Whitney U and Friedman tests.


Results: ChatGPT-4o achieved the highest overall accuracy across both Pediatric (90.88%) and Restorative Dentistry (93.02%), whereas ChatGPT-3.5 exhibited the lowest performance in Pediatric Dentistry (70.71%). Statistically significant inter-model differences were observed in Pediatric Dentistry (p = 0.002), but not in Restorative Dentistry (p = 0.254). Temporal analyses highlighted the consistent performance of ChatGPT-4o, while Gemini 2.0 Flash displayed marked fluctuations in accuracy across certain years and subdomains.


Conclusion: Among the evaluated LLMs, ChatGPT-4o demonstrated superior and stable performance in addressing dental specialty exam questions, particularly in content requiring contextual comprehension. These findings underscore the potential of advanced AI tools to enhance dental education and examination readiness. Nonetheless, appropriate model selection should be tailored to the complexity of the task, domain specificity, and pedagogical objectives.

Full text article

Generated from XML file

Authors

Ayşenur Altuğ Yıldırım
aysenuraltug@gmail.com (Primary Contact)
Menekşe Alim
1.
Altuğ Yıldırım A, Alim M. Evaluating the accuracy of ChatGPT and Gemini models in dental specialty exams: A comparative study of pediatric and restorative dentistry. Int Dent Res. 2026;16(S1):e260673. doi:10.5577/intdentres.673

Article Details

How to Cite

1.
Altuğ Yıldırım A, Alim M. Evaluating the accuracy of ChatGPT and Gemini models in dental specialty exams: A comparative study of pediatric and restorative dentistry. Int Dent Res. 2026;16(S1):e260673. doi:10.5577/intdentres.673
Smart Citations via scite_

Similar Articles

You may also start an advanced similarity search for this article.

The Effect of Clinical Education on Dental Students’ Stress and Anxiety Levels

Lect. Dr. Seyma Eken, Assoc. Prof. Berceste Guler Ayyildiz, Prof. Fitnat Deniz Cetiner, DDS Anıl...
Abstract View : 0

Accuracy comparison of chatbot responses in temporomandibular joint disorders

Esra Nur Avukat, Mirac Berke Topcu Ersöz , Canan Akay
Abstract View : 0

Accuracy comparison of chatbot responses in temporomandibular joint disorders

Esra Nur Avukat, Mirac Berke Topcu Ersöz , Canan Akay
Abstract View : 0