arXiv:2511.22055cs.CVcs.MM2025-11被引 11

首个专用于牙科的多模态大模型,提升影像分析与诊断可信度。

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

  • 构建临床思维链数据集,引导模型模仿牙医诊断逻辑。
  • 在牙科多模态基准上得分51.84,显著超越GPT-5。
  • 适合牙科AI研究者与医疗智能化开发者使用。

多模态大语言模型在多个医学领域展现出巨大潜力,但牙科仍缺乏深入探索,受限于领域专用数据少、专家标注稀缺、模态建模不足及可靠性挑战。本文提出OralGPT-Omni,首个面向牙科的专用多模态大模型,支持多种牙科影像模态与临床任务的全面可信分析。为捕捉牙医诊断推理过程,我们构建了TRACE-CoT——一个基于临床的思维链数据集。结合四阶段训练范式,显著增强模型对牙科图像的理解能力。同时,提出MMOral-Uni,首个统一的牙科多模态分析基准,包含2,809个开放问答对,覆盖五种模态和五类任务。OralGPT-Omni在该基准上取得51.84分,在MMOral-OPG上得45.31分,远超GPT-5表现。研究成果推动智能牙科发展,代码、模型与数据集将公开共享。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. In this paper, we present OralGPT-Omni, the first dental-specialized MLLM designed for comprehensive and trustworthy analysis across diverse dental imaging modalities and clinical tasks. To explicitly capture dentists' diagnostic reasoning, we construct TRACE-CoT, a clinically grounded chain-of-thought dataset that mirrors dental radiologists' decision-making processes. This reasoning supervision, combined with our proposed four-stage training paradigm, substantially strengthens the model's capacity for dental image understanding and analysis. In parallel, we introduce MMOral-Uni, the first unified multimodal benchmark for dental image analysis. It comprises 2,809 open-ended question-answer pairs spanning five modalities and five tasks, offering a comprehensive evaluation suite to date for MLLMs in digital dentistry. OralGPT-Omni achieves an overall score of 51.84 on the MMOral-Uni benchmark and 45.31 on the MMOral-OPG benchmark, dramatically outperforming the scores of GPT-5. Our work promotes intelligent dentistry and paves the way for future advances in dental image analysis. All code, benchmark, and models will be made publicly available.

牙科AI多模态模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。