用DINOv3统一处理牙科影像,无需微调就能高效分析全景片和口内照片。
DinoDental: Benchmarking DINOv3 as a Unified Vision Encoder for Dental Image Analysis
- 用自监督大模型DINOv3作为通用视觉编码器,直接迁移至牙科影像任务。
- 在全景片和口内照片上均表现良好,尤其在边界敏感任务中优势明显。
- 适合缺乏标注数据的牙科AI研究者快速部署模型,省去预训练成本。
牙科影像标注稀缺且成本高昂,制约了人工智能在该领域的应用。本文提出DinoDental,一个统一基准,系统评估DINOv3(在17亿张图像上预训练的自监督视觉基础模型)是否可作为无需领域预训练的可靠编码器,用于牙科图像分析。该基准涵盖多个公开数据集,覆盖全景牙片与口内照片的分类、检测、实例分割等任务。通过调整模型规模、输入分辨率,并对比冻结特征、全微调和低秩适应(LoRA)等策略,实验表明DINOv3在两类影像上均具强泛化能力,尤其在口内照片理解和边界敏感密集预测任务中表现突出。DinoDental为牙科AI提供了系统评估框架,推动高效模型选择与适配。
原文摘要 · Abstract (English)
The scarcity and high cost of expert annotations in dental imaging present a significant challenge for the development of AI in dentistry. DINOv3, a state-of-the-art, self-supervised vision foundation model pre-trained on 1.7 billion images, offers a promising pathway to mitigate this issue. However, its reliability when transferred to the dental domain, with its unique imaging characteristics and clinical subtleties, remains unclear. To address this, we introduce DinoDental, a unified benchmark designed to systematically evaluate whether DINOv3 can serve as a reliable, off-the-shelf encoder for comprehensive dental image analysis without requiring domain-specific pre-training. Constructed from multiple public datasets, DinoDental covers a wide range of tasks, including classification, detection, and instance segmentation on both panoramic radiographs and intraoral photographs. We further analyze the model's transfer performance by scaling its size and input resolution, and by comparing different adaptation strategies, including frozen features, full fine-tuning, and the parameter-efficient Low-Rank Adaptation (LoRA) method. Our experiments show that DINOv3 can serve as a strong unified encoder for dental image analysis across both panoramic radiographs and intraoral photographs, remaining competitive across tasks while showing particularly clear advantages for intraoral image understanding and boundary-sensitive dense prediction. Collectively, DinoDental provides a systematic framework for comprehensively evaluating DINOv3 in dental analysis, establishing a foundational benchmark to guide efficient and effective model selection and adaptation for the dental AI community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。