首个多模态牙科分诊基准,助力临床级AI决策系统研发
Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

- 构建真实门诊流程的专家标注数据集,含246例病例与完整推理路径
- 模型在细粒度治疗分诊上与初级牙医差距显著,错误集中于多领域病例
- 揭示需同时依赖主诉和全景片信息,适合临床AI安全验证研究者使用
牙科分诊是一项关乎安全的临床分流任务,需整合患者主诉与影像证据等多模态信息以制定完整转诊方案。本文提出Dental-TriageBench,首个由专家标注的、面向推理驱动的多模态牙科分诊基准。该数据集基于真实门诊流程构建,包含246例去标识化病例,每例均配有专家撰写的黄金推理轨迹及层级化分诊标签。我们对19个专有、开源及医疗领域多模态大模型(MLLMs)进行评测,以三位初级牙医作为人类基线,发现模型在细粒度治疗级别分诊上存在显著人机差距。进一步分析表明,准确分诊需同时依赖主诉与全口全景片(OPG)信息,且模型错误集中在多转诊领域病例,常表现为推荐范围过窄与遗漏型错误。Dental-TriageBench为开发更符合临床实际、覆盖全面且安全的多模态医疗AI系统提供了真实测试平台。
原文摘要 · Abstract (English)
Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiographic evidence) to determine complete referral plans. We present Dental-TriageBench, the first expert-annotated benchmark for reasoning-driven multimodal dental triage. Built from authentic outpatient workflows, it contains 246 de-identified cases annotated with expert-authored golden reasoning trajectories, together with hierarchical triage labels. We benchmark 19 proprietary, open-source, and medical-domain MLLMs against three junior dentists serving as the human baseline, and find a substantial human--model gap, on fine-grained treatment-level triage. Further analyses show that accurate triage requires both complaint and OPG information, and that model errors concentrate on cases with multiple referral domains, where MLLMs tend to produce overly narrow referral sets and omission-heavy errors. Dental-TriageBench provides a realistic testbed for developing multimodal clinical AI systems that are more clinically grounded, coverage-aware, and safer for downstream care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。