构建牙科全景片细粒度标注基准,评估大模型诊断一致性。
PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

- 基于专家标注的9524张牙片,实现牙齿级结构-语言一致标注。
- 现有大模型报告流畅但临床错误多,定位与诊断准确率低。
- 适合牙科AI研究者、医学大模型评测人员使用。
当前多模态大语言模型在牙科全景片(OPG)中的评估受限于缺乏细粒度、临床可靠的基准。本文提出PanDent,一个大规模、基于临床的OPG基准,包含9,524张高质量全景片,每张均由经验牙医生成结构化标注,并经口腔颌面放射科医生验证,提供可靠的牙齿级诊断监督。基于临床定义的报告逻辑,将专家验证发现转化为临床一致的放射报告,建立结构化证据与自由文本描述之间的明确对应关系。该设计可评估大模型报告是否不仅语言通顺,且与专家验证的牙齿级结果临床一致。在多种大模型(包括SOTA专有模型、通用开源模型和医疗专用模型)上测试显示,当前模型虽能生成流畅报告,但在细粒度定位和牙齿级诊断上存在显著错误。在PanDent上微调后,模型的结构-语言一致性显著提升,视觉定位准确率与诊断正确率明显改善,输出更接近专家水平。本研究确立了潘登作为牙科大模型牙齿级临床推理评估的严格基准,为临床导向牙科AI提供了宝贵资源。
原文摘要 · Abstract (English)
Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-grained, clinically reliable benchmarks that reflect expert interpretation. This work introduces PanDent, a large-scale, clinically grounded OPG benchmark built upon fine-grained, expert-validated tooth-level annotations. The dataset comprises 9,524 high-quality OPGs, each associated with comprehensive structured annotations produced by experienced dentists and further validated by an oral and maxillofacial radiologist, providing clinically reliable supervision for tooth-level diagnosis and reasoning. Clinically consistent radiology reports are constructed from expert-validated findings using clinician-defined reporting logic, establishing explicit correspondence between structured clinical evidence and free-text descriptions. This design enables evaluation of whether MLLMs generate reports that are not only linguistically coherent but also clinically consistent with expert-validated tooth-level findings. Experiments are conducted on diverse MLLMs, including state-of-the-art (SOTA) proprietary models, general-domain open-source models, and medical-specific models. Results show that current MLLMs can generate fluent reports, yet fail to produce clinically consistent descriptions, exhibiting substantial errors in fine-grained localization and tooth-level diagnosis. Fine-tuning on PanDent significantly improves structure-language consistency, substantially enhancing visual localization accuracy and diagnostic correctness, and bringing model outputs closer to expert dental interpretation. These results establish PanDent as a rigorous benchmark for evaluating tooth-level clinical reasoning in MLLMs and a valuable resource for clinically grounded dental AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。