首个跨国牙科大模型推理评估基准,揭示当前模型在复杂临床决策中严重不安全。
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

- 构建覆盖88国14个牙科领域的多格式专家验证题库,分三层推理能力评估
- 模型在案例推理题上准确率仅22.34%,且31.01%的建议存在安全隐患
- 特别暴露正畸等专科风险,为医疗AI安全落地提供关键评估工具
尽管大型语言模型(LLMs)在医学领域具有变革潜力,但其在真实临床场景中的推理鲁棒性和安全性仍严重缺乏研究,尤其在牙科领域。本文提出GlobalDentBench,首个跨国牙科基准,涵盖6大洲88个国家和地区的14个牙科专业。该基准包含8,978道专家验证题目,分为选择题、简答题和病例题三种形式,评估知识回忆(L1)、常规推理(L2)和个性化推理(L3)三个层级。通过六位资深牙医校准,自动化构建框架在选择题与简答题上达成99.98%专家一致性,在更复杂的病例题上达96.78%。对12个前沿大模型的评估显示,随着推理复杂度提升,性能显著下降:选择题准确率为81.34%,简答题降至64.53%,病例题仅22.34%;在三层次中分别从74.01%、55.64%降至35.71%。更重要的是,真实病例风险分析发现,大模型生成建议的整体不安全率达31.01%,其中4.51%可能导致不可逆患者伤害,且正畸等专科尤为突出。这些结果揭示了当前大模型在医学推理与安全方面的根本局限。GlobalDentBench为可信临床AI评估提供了可扩展基础,强调在医疗部署前必须进行严格验证。
原文摘要 · Abstract (English)
While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。