arXiv:2512.11558cs.CVcs.AI2025-12被引 6

用高质量牙科数据训练,让大模型更懂牙齿细节和诊断推理。

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

  • 注入专业牙科知识,分阶段训练提升多模态理解
  • 在12万张牙片上训练,70亿参数模型表现超主流水平
  • 适合医疗AI研究者、牙科数字化从业者参考

可靠的牙科多模态数据解读对自动化口腔健康至关重要,但现有多模态大模型难以捕捉精细的视觉特征,且推理能力不足。为此,我们提出DentalGPT,一种通过高质量领域知识注入与强化学习构建的专用牙科多模态大模型。我们整合超过12万张牙科图像及其详细描述,构建了迄今最大的牙科多模态标注数据集,涵盖最具诊断意义的视觉特征。基于该数据集训练显著提升了模型对牙病的视觉理解,后续强化学习进一步增强其多模态复杂推理能力。在口内及全景影像基准测试,以及医学VQA基准的牙科子集上全面评估显示,DentalGPT在疾病分类与牙科VQA任务中表现优异,尽管仅含70亿参数,仍优于多个先进MLLM。结果表明,高质量牙科数据结合分阶段适配,是打造高效专用牙科多模态大模型的有效路径。

原文摘要 · Abstract (English)

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack sufficient reasoning ability for precise diagnosis. To address these limitations, we present DentalGPT, a specialized dental MLLM developed through high-quality domain knowledge injection and reinforcement learning. Specifically, the largest annotated multimodal dataset for dentistry to date was constructed by aggregating over 120k dental images paired with detailed descriptions that highlight diagnostically relevant visual features, making it the multimodal dataset with the most extensive collection of dental images to date. Training on this dataset significantly enhances the MLLM's visual understanding of dental conditions, while the subsequent reinforcement learning stage further strengthens its capability for multimodal complex reasoning. Comprehensive evaluations on intraoral and panoramic benchmarks, along with dental subsets of medical VQA benchmarks, show that DentalGPT achieves superior performance in disease classification and dental VQA tasks, outperforming many state-of-the-art MLLMs despite having only 7B parameters. These results demonstrate that high-quality dental data combined with staged adaptation provides an effective pathway for building capable and domain-specialized dental MLLMs.

牙科AI多模态大模型医学视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。