We are pleased to announce that our latest research has been published in Scientific Reports (Nature Portfolio, Open Access).
The study, titled “Benchmarking GPT-5, Gemini 2.5 Pro, Grok 4, and other LLMs on pediatric dentistry questions from a dental specialization exam,” evaluates 11 recent large language models (17 configurations including reasoning modes) on 119 pediatric dentistry questions from the last ten years of Türkiye’s Dentistry Specialization Examination (DUS). Accuracy and response-generation time were measured and statistically compared for each model.
Gemini 2.5 Pro (92.44%) and GPT-5 (90.76%) approached expert-level accuracy, followed by Grok-4 (88.24%), while Qwen-3 and MedGemma performed markedly worse. A clear speed–accuracy trade-off emerged: reasoning-intensive modes such as DeepSeek R1 improved scores but required up to 68 s per answer, whereas fast models (<1 s) were less accurate. The results underline the need for careful validation before LLMs are used in high-stakes dental education or assessment.
Authors:
Şükriye Türkoğlu Kayacı (University of Health Sciences)
Hamza Osman İlhan (Yıldız Technical University)
Melek Taşsöker (Necmettin Erbakan University)
Taha Çap
Helin Demir (University of Health Sciences)




