arXiv:2510.14532cs.CV2025-10被引 2

首个面向口腔颌面放射学的通用视觉模型,提升AI诊断泛化能力。

Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology

  • 基于自监督学习构建多模态牙科影像基础模型,支持多种任务
  • 在8个亚专科、160万张图像上验证,性能超越主流方法
  • 适合需要高效、跨模态诊断的临床与研究场景

口腔颌面放射学在牙科医疗中至关重要,但影像解读受限于专业人才短缺。现有AI系统因单模态、任务专一及依赖昂贵标注数据,难以泛化至多样临床场景。为此,我们提出DentVFM,首个专为牙科设计的视觉基础模型家族。DentVFM基于约160万张来自多个医疗机构的多模态牙科影像构建的DentVista数据集,采用自监督学习,包含基于ViT架构的2D与3D变体。为填补牙科智能评估空白,我们引入DentBench基准,覆盖8个牙科亚专科、更多疾病类型、多种成像模态及广泛地理分布。实验表明,DentVFM在疾病诊断、治疗分析、生物标志物识别及解剖标志点检测分割等任务中表现出色,显著优于监督、自监督和弱监督基线,具备更强泛化性、标签效率与可扩展性。此外,其跨模态诊断能力在传统影像缺失时表现优于资深牙医。DentVFM为牙科AI树立新范式,推动智能牙科医疗发展,缓解全球口腔健康服务缺口。

原文摘要 · Abstract (English)

Oral and maxillofacial radiology plays a vital role in dental healthcare, but radiographic image interpretation is limited by a shortage of trained professionals. While AI approaches have shown promise, existing dental AI systems are restricted by their single-modality focus, task-specific design, and reliance on costly labeled data, hindering their generalization across diverse clinical scenarios. To address these challenges, we introduce DentVFM, the first family of vision foundation models (VFMs) designed for dentistry. DentVFM generates task-agnostic visual representations for a wide range of dental applications and uses self-supervised learning on DentVista, a large curated dental imaging dataset with approximately 1.6 million multi-modal radiographic images from various medical centers. DentVFM includes 2D and 3D variants based on the Vision Transformer (ViT) architecture. To address gaps in dental intelligence assessment and benchmarks, we introduce DentBench, a comprehensive benchmark covering eight dental subspecialties, more diseases, imaging modalities, and a wide geographical distribution. DentVFM shows impressive generalist intelligence, demonstrating robust generalization to diverse dental tasks, such as disease diagnosis, treatment analysis, biomarker identification, and anatomical landmark detection and segmentation. Experimental results indicate DentVFM significantly outperforms supervised, self-supervised, and weakly supervised baselines, offering superior generalization, label efficiency, and scalability. Additionally, DentVFM enables cross-modality diagnostics, providing more reliable results than experienced dentists in situations where conventional imaging is unavailable. DentVFM sets a new paradigm for dental AI, offering a scalable, adaptable, and label-efficient model to improve intelligent dental healthcare and address critical gaps in global oral healthcare.

牙科AI视觉模型自监督学习通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。