arXiv:2512.19213cs.CV2025-12

用生成图像替代真实数据重放,实现医疗多模态影像持续自监督学习。

InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training

  • 通过模型反演生成合成图像,避免访问历史真实数据。
  • 在9个下游任务中表现媲美或优于依赖数据重放的方法。
  • 适合数据隐私要求高、跨机构协作受限的医学影像场景。

医疗影像中的持续自监督学习(CSSL)通过顺序训练基础模型,减少对多模态图像联合收集的需求,提升下游任务性能并保护数据隐私。然而,现有方法仍依赖回放历史数据以防止灾难性遗忘,损害隐私且限制实际应用。本文提出InvCoSS,一种基于反演的持续自监督学习框架。训练完前一任务后,该方法反演预训练模型生成逼近原始分布的合成图像,与新任务数据联合优化,有效缓解遗忘问题,同时严格满足不访问真实历史数据的要求。为提升合成图像保真度,引入具有多尺度融合结构的InvUNet,恢复图像的高低频成分;为增强多样性并防止模式崩溃,设计无类别引导的排斥表示学习机制,促进合成图像特征空间多样化。在9个下游任务上的实验表明,InvCoSS性能可媲美甚至超越传统数据重放方法,显著降低存储开销,消除数据隐私顾虑。

原文摘要 · Abstract (English)

Continual self-supervised learning (CSSL) in medical imaging trains a foundation model sequentially, alleviating the need for collecting multi-modal images for joint training and offering promising improvements in downstream performance while preserving data privacy. However, most existing methods still rely on replaying data from previous stages to prevent catastrophic forgetting, which compromises privacy and limits their applicability in real-world scenarios where data transfer across sites is often restricted. In this work, we propose InvCoSS, an inversion-driven continual self-supervised learning framework for medical multi-modal image pre-training. Specifically, after training on a previous task, InvCoSS inverts the pre-trained self-supervised model to generate synthetic images that approximate the original training distribution. These synthetic images are then combined with data from the new task for joint optimization, which effectively mitigates catastrophic forgetting while strictly adhering to the constraint of no access to previous real data. Furthermore, to improve the fidelity of synthetic images, we introduce a novel InvUNet with a multi-scale fusion architecture to restore both high- and low-frequency components of the inverted images. To enhance diversity and prevent mode collapse, we design a repulsive representation-learning mechanism that encourages a diverse feature space for synthetic images without class guidance. Extensive experiments across nine downstream tasks validate the effectiveness of InvCoSS, achieving performance comparable to or even superior to prior data-replay methods while significantly reducing storage requirements and eliminating data privacy constraints.

医疗影像持续学习自监督图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。