arXiv:2605.01333cs.CL2026-05

构建牙科影像认知能力评测基准,揭示大模型与医生的差距

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

  • 设计涵盖四类认知能力的牙科影像评测任务
  • 基于3820份临床评估,对比6个前沿大模型表现
  • 针对牙科诊疗流程提供可落地的改进方向

多模态大语言模型在牙科影像分析中展现出潜力,但其对放射影像分析所需多层次认知过程的理解仍不明确。本文提出一个全面的评测基准,覆盖根尖片、全景片和侧头颅片三种关键影像模态,定义了感知、理解、预测和决策四类认知能力。基准包含27项源自公开数据集的临床任务,配有手工标注和3820次临床医师评估。评估了GPT-5.2、GLM-4.6等六种前沿多模态大模型,揭示了模型与临床医生之间的性能差距,分析了模型优劣与失败模式,并提出改进建议。该数据资源将推动下一代符合临床认知、安全要求和工作流复杂性的AI系统发展。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-level cognitive processes required for radiographic analysis remains unclear. Here, we present a comprehensive benchmark to evaluate the cognitive capabilities of MLLMs in dental radiographic analysis. It spans three critical imaging modalities, i.e., periapical, panoramic, and lateral cephalometric radiographs, and defines four cognitive categories: perception, comprehension, prediction, and decision-making. The benchmark comprises 27 clinically grounded tasks derived from public datasets, with manually curated annotations and 3,820 clinician assessments for evaluation. Six frontier MLLMs, including GPT-5.2 and GLM-4.6, are evaluated. We demonstrate the performance gap between MLLMs and clinicians in dental practice, delineate model strengths and limitations, characterize failure patterns, and provide recommendations for improvement. This data resource will facilitate the development of next-generation artificial intelligence systems aligned with clinical cognition, safety requirements, and workflow complexity in dental practice.

多模态模型牙科AI认知评测医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。