首个面向全景牙片分析的多模态数据集与评测基准,推动口腔AI发展。
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- 构建包含1.3万条指令的全景牙片多模态数据集,覆盖多种诊断任务。
- 64个大模型在新评测中最高仅达41.45%准确率,暴露当前模型短板。
- 基于该数据集微调的OralGPT模型提升24.73%,助力临床应用。
近年来大型视觉语言模型(LVLMs)在通用医疗任务中表现优异,但在牙科等专业领域仍缺乏深入研究。全景牙片因其结构密集、病灶细微,现有医学基准和指令数据集难以覆盖其复杂性。为此,我们提出MMOral,首个专为全景牙片解读设计的大规模多模态指令数据集与评测基准。该数据集包含20,563张标注图像及130万条指令实例,涵盖属性提取、报告生成、视觉问答和图像对话等任务。同时,我们构建了涵盖五大诊断维度的MMOral-Bench评测体系。在该基准上评估64个LVLMs发现,即使最优模型GPT-4o也仅达到41.45%准确率,揭示当前模型在该领域的显著局限。为推进发展,我们提出OralGPT,基于Qwen2.5-VL-7B进行监督微调(SFT),利用精心构建的MMOral指令数据集。单轮微调即带来显著性能提升,例如OralGPT实现24.73%的准确率增益。MMOral与OralGPT有望成为智能牙科的重要基础,推动更具临床价值的多模态系统落地。数据集、模型、评测工具均已开源。
原文摘要 · Abstract (English)
Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular, panoramic X-rays, a widely used imaging modality in oral radiology, pose interpretative challenges due to dense anatomical structures and subtle pathological cues, which are not captured by existing medical benchmarks or instruction datasets. To this end, we introduce MMOral, the first large-scale multimodal instruction dataset and benchmark tailored for panoramic X-ray interpretation. MMOral consists of 20,563 annotated images paired with 1.3 million instruction-following instances across diverse task types, including attribute extraction, report generation, visual question answering, and image-grounded dialogue. In addition, we present MMOral-Bench, a comprehensive evaluation suite covering five key diagnostic dimensions in dentistry. We evaluate 64 LVLMs on MMOral-Bench and find that even the best-performing model, i.e., GPT-4o, only achieves 41.45% accuracy, revealing significant limitations of current models in this domain. To promote the progress of this specific domain, we also propose OralGPT, which conducts supervised fine-tuning (SFT) upon Qwen2.5-VL-7B with our meticulously curated MMOral instruction dataset. Remarkably, a single epoch of SFT yields substantial performance enhancements for LVLMs, e.g., OralGPT demonstrates a 24.73% improvement. Both MMOral and OralGPT hold significant potential as a critical foundation for intelligent dentistry and enable more clinically impactful multimodal AI systems in the dental field. The dataset, model, benchmark, and evaluation suite are available at https://github.com/isbrycee/OralGPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。