用离散自监督让医疗影像模型更通用,少调参也能用。
DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- 通过多尺度离散量化构建特征瓶颈,强制模型学结构化特征。
- 在分类与分割任务中表现优异,低标签数据下仍高效。
- 适合医疗影像迁移学习,尤其对标注少的场景友好。
自监督学习(SSL)已成为医疗图像表征学习的重要范式,尤其在标注数据有限的场景下。然而,现有方法常依赖复杂架构、解剖先验或高度调优的增强策略,限制了其可扩展性和泛化能力。更关键的是,在胸片等解剖结构相似、病灶细微的模态中,模型易陷入捷径学习。本文提出DiSSECT——一种基于离散自监督的高效临床可迁移表征框架,将多尺度向量量化引入SSL流程,形成离散表征瓶颈,迫使模型学习重复性强、结构感知的特征,抑制视图特异性或低效模式。该方法在分类与分割任务上均取得优异表现,几乎无需微调,并在低标签条件下展现出极高标签效率。我们在多个公开医疗影像数据集上验证了DiSSECT,结果表明其鲁棒性与泛化能力优于当前主流方法。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has emerged as a powerful paradigm for medical image representation learning, particularly in settings with limited labeled data. However, existing SSL methods often rely on complex architectures, anatomy-specific priors, or heavily tuned augmentations, which limit their scalability and generalizability. More critically, these models are prone to shortcut learning, especially in modalities like chest X-rays, where anatomical similarity is high and pathology is subtle. In this work, we introduce DiSSECT -- Discrete Self-Supervision for Efficient Clinical Transferable Representations, a framework that integrates multi-scale vector quantization into the SSL pipeline to impose a discrete representational bottleneck. This constrains the model to learn repeatable, structure-aware features while suppressing view-specific or low-utility patterns, improving representation transfer across tasks and domains. DiSSECT achieves strong performance on both classification and segmentation tasks, requiring minimal or no fine-tuning, and shows particularly high label efficiency in low-label regimes. We validate DiSSECT across multiple public medical imaging datasets, demonstrating its robustness and generalizability compared to existing state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。