Evaluating the accuracy of ChatGPT and Gemini models in dental specialty exams: A comparative study of pediatric and restorative dentistry
Abstract
Aim: The integration of large language models (LLMs) such as ChatGPT and Gemini into healthcare education presents novel opportunities for knowledge dissemination and assessment. Despite their growing prominence, empirical evidence regarding their reliability and domain-specific accuracy in dental specialty examinations remains limited. This study aimed to evaluate and compare the performance of four AI-driven LLMs—ChatGPT-3.5, ChatGPT-4o, Gemini 2.0 Flash, and Gemini 2.0 Advanced— by employing official past questions from the Turkish Dental Specialty Examination (DUS) in Pediatric Dentistry and Restorative Dentistry.
Methodology: A total of 178 multiple-choice questions (90 in Restorative Dentistry and 88 in Pediatric Dentistry) were obtained from DUS examinations conducted between 2012 and 2017. All items were presented in Turkish to each model via standardized prompts in isolated sessions to prevent cross-interaction and memory bias. Response accuracy was determined by evaluating the correctness of the chosen answers. Comparative analyses between models, domains, and annual trends were conducted using non-parametric statistical tests, including the Mann–Whitney U and Friedman tests.
Results: ChatGPT-4o achieved the highest overall accuracy across both Pediatric (90.88%) and Restorative Dentistry (93.02%), whereas ChatGPT-3.5 exhibited the lowest performance in Pediatric Dentistry (70.71%). Statistically significant inter-model differences were observed in Pediatric Dentistry (p = 0.002), but not in Restorative Dentistry (p = 0.254). Temporal analyses highlighted the consistent performance of ChatGPT-4o, while Gemini 2.0 Flash displayed marked fluctuations in accuracy across certain years and subdomains.
Conclusion: Among the evaluated LLMs, ChatGPT-4o demonstrated superior and stable performance in addressing dental specialty exam questions, particularly in content requiring contextual comprehension. These findings underscore the potential of advanced AI tools to enhance dental education and examination readiness. Nonetheless, appropriate model selection should be tailored to the complexity of the task, domain specificity, and pedagogical objectives.
Full text article
Authors
Copyright © 2026 The Author(s)

This work is licensed under a Creative Commons Attribution 4.0 International License.
This is an Open Access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.