arXiv:2505.15248cs.CVcs.LG2025-05被引 2

利用多视角兽医影像自监督学习,提升模型对解剖结构的理解能力。

VET-DINO: Learning Anatomical Understanding Through Multi-View Distillation in Veterinary Imaging

  • 基于同一病例的多视角影像进行知识蒸馏,学习视图不变的解剖特征。
  • 在500万张犬类放射影像上训练,显著优于纯合成数据增强方法。
  • 适用于兽医影像分析,推动医学自监督学习向领域特性靠拢。

自监督学习已成为训练深度神经网络的强大范式,尤其在标注数据稀缺的医学影像领域。现有方法通常依赖单图的合成增强,而本文提出VET-DINO框架,利用医学影像的独特优势:同一病例可获取多个标准化视角。通过一系列临床兽医放射影像,模型得以学习视图不变的解剖结构,并从2D投影中隐含地建立3D理解。我们在包含66.8万例犬类研究、总计500万张影像的数据集上验证该方法。大量实验(包括视图合成与下游任务性能)表明,使用真实多视角配对数据的学习效果显著优于纯合成增强。VET-DINO在多种兽医影像任务中达到当前最优性能。本工作确立了医学自监督学习的新范式,强调利用领域特有属性,而非简单套用自然图像技术。

原文摘要 · Abstract (English)

Self-supervised learning has emerged as a powerful paradigm for training deep neural networks, particularly in medical imaging where labeled data is scarce. While current approaches typically rely on synthetic augmentations of single images, we propose VET-DINO, a framework that leverages a unique characteristic of medical imaging: the availability of multiple standardized views from the same study. Using a series of clinical veterinary radiographs from the same patient study, we enable models to learn view-invariant anatomical structures and develop an implied 3D understanding from 2D projections. We demonstrate our approach on a dataset of 5 million veterinary radiographs from 668,000 canine studies. Through extensive experimentation, including view synthesis and downstream task performance, we show that learning from real multi-view pairs leads to superior anatomical understanding compared to purely synthetic augmentations. VET-DINO achieves state-of-the-art performance on various veterinary imaging tasks. Our work establishes a new paradigm for self-supervised learning in medical imaging that leverages domain-specific properties rather than merely adapting natural image techniques.

自监督学习兽医影像多视角学习解剖理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。