arXiv:2605.23995cs.CVcs.AI2026-05综述

医学影像自监督学习需匹配任务需求,否则会适得其反。

Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines

论文配图:Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
图 1 · 摘自论文原文
  • 按下游任务设计自监督目标,避免信息错配
  • 对比类方法适合分类但忽略局部病灶,重建类更利于分割
  • 低标注场景下尤需注意数据增强与评估指标的临床相关性

自监督学习(SSL)在医学影像分析中日益重要,通过无标注数据学习可迁移表征以减少对昂贵专家标注的依赖。然而,SSL性能不仅取决于模型架构,更取决于自监督目标是否保留下游临床任务所需信息。本综述分析了2017至2025年发表的78项研究,将其归纳为四类范式:对比型、非对比与预测型、生成与重建型、混合型。不按时间顺序罗列,而是考察各类范式在分类、分割、检测、重建和回归任务中的适用性。研究表明,有效性由目标、模态与任务之间的匹配程度决定,而非单一策略。对比目标擅长全局判别表征,适用于分类,但可能弱化局部病变;空间预测、掩码建模与重建目标更利于保留解剖结构,适合分割与密集预测。关键问题是,目标不匹配会导致负向迁移,如通过成像特征或增强手段学习捷径,反而抹除诊断信号。在低标签情况下SSL最有效,但其效果依赖模态感知的增强方式、病理信息保持的扰动设计及临床意义的评估。最后提出实用设计指南与开放挑战。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data. However, SSL performance depends not only on model architecture but also on whether the self-supervised objective preserves the information required by the downstream clinical task. This review presents a task-oriented synthesis of SSL methods for medical imaging, focusing on how the design of the self-supervised objective interacts with imaging modality, label availability, and downstream performance. We analyze $78$ studies published from 2017 to 2025 and organize them into four paradigms: contrastive, non-contrastive and predictive, generative and reconstruction-based, and hybrid learning. Rather than cataloging methods chronologically, we examine how these paradigms support classification, segmentation, detection, reconstruction, and regression. The evidence suggests that effectiveness is governed by the match among objective, modality, and downstream task rather than by any single strategy. Contrastive objectives favor global discriminative representations suited to classification but may underrepresent localized pathology, whereas spatial-prediction, masked-modeling, and reconstruction objectives better preserve anatomical structure for segmentation and dense prediction. Critically, misaligned objectives can cause negative transfer through shortcut learning on acquisition signatures or augmentation that erases diagnostic signal rather than merely weaker gains. SSL is most beneficial in low-label regimes, but its effectiveness depends on modality-aware augmentation, pathology-preserving corruption, and clinically meaningful evaluation. We conclude with practical design guidelines and open challenges for clinically aligned SSL.

自监督学习医学影像任务对齐临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。