arXiv:2509.06467cs.CV2025-09被引 10

DINOv3无需微调即可在医学影像中表现优异,但特定场景下仍有局限。

Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration

  • 用自然图像预训练的DINOv3直接用于医学任务,无需领域微调。
  • 在多种2D/3D任务中超越生物医学专用模型,部分任务达新基准。
  • 全切片图像等深度专业化场景性能下降,且不遵循典型缩放规律。

大规模视觉基础模型在自然图像上预训练后,正推动计算机视觉范式变革。然而,这些前沿模型在医学影像等专业领域的迁移能力仍不明确。本研究评估DINOv3——一种基于自然图像自监督训练的视觉变换器(ViT),是否可作为无需领域微调的通用医学视觉编码器。我们在多种医学影像模态上系统测试其在2D与3D分类、分割和配准任务中的表现,并分析不同模型规模与输入分辨率下的可扩展性。结果表明,DINOv3展现出卓越性能,确立了强有力的新型基准。令人惊讶的是,它在多个任务中甚至优于仅在医学数据上训练的BiomedCLIP和CT-Net。然而,我们发现其在需要深度领域特化的场景中存在明显短板,如全切片图像(WSI)、电子显微镜(EM)和正电子发射断层扫描(PET)。此外,DINOv3在医学领域内不一致地遵循缩放规律:模型越大或特征分辨率越高,性能并非始终提升,不同任务间表现出各异的缩放行为。总体而言,本工作确立了DINOv3作为强基准,其强大视觉特征可作为多类医学任务的稳健先验,未来可用于增强3D重建中的多视角一致性。

原文摘要 · Abstract (English)

The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundation models' efficacies transfer to specialised domains such as medical imaging remains an open question. This report investigates whether DINOv3, a state-of-the-art self-supervised vision transformer (ViT) pre-trained on natural images, can directly serve as a powerful, unified encoder for medical vision tasks without domain-specific fine-tuning. To answer this, we benchmark DINOv3 across common medical vision tasks, including 2D and 3D classification, segmentation, and registration on a wide range of medical imaging modalities. We systematically analyse its scalability by varying model sizes and input image resolutions. Our findings reveal that DINOv3 shows impressive performance and establishes a formidable new baseline. Remarkably, it can even outperform medical-specific foundation models like BiomedCLIP and CT-Net on several tasks, despite being trained solely on natural images. However, we identify clear limitations: The model's features degrade in scenarios requiring deep domain specialisation, such as in whole-slide images (WSIs), electron microscopy (EM), and positron emission tomography (PET). Furthermore, we observe that DINOv3 does not consistently follow the scaling law in the medical domain. Its performance does not reliably increase with larger models or finer feature resolutions, showing diverse scaling behaviours across tasks. Overall, our work establishes DINOv3 as a strong baseline, whose powerful visual features can serve as a robust prior for multiple medical tasks. This opens promising future directions, such as leveraging its features to enforce multiview consistency in 3D reconstruction.

医学影像视觉基础模型自监督学习缩放规律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。