让AI更懂影像细节,提升骨骼病变识别精度。
AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images
- 基于解剖结构引导视觉变压器的对比学习
- 在少标注下实现多标签分类超越7种方法
- 仅用图像级标签就能精准定位异常区域
现有自监督方法如对比学习主要关注全局判别,忽略了放射影像分析所需的精细解剖细节。为此,我们提出一种解剖驱动的自监督框架AFiRe,以增强放射影像中的细粒度表征。AFiRe通过两项自监督策略协同工作:(i) 逐标记解剖引导对比学习,依据结构与类别一致性对图像标记进行对齐,提升细粒度空间-解剖判别力;(ii) 像素级异常去除重建,特别关注局部异常,从而以详细几何信息优化学习判别能力。此外,我们提出合成病灶掩码,在保留内部一致性的同时增强解剖多样性,避免传统数据增强(如裁剪、仿射变换)造成的破坏。实验表明,AFiRe:(i) 提供稳健的解剖判别能力,相比最先进对比学习方法生成更凝聚的特征簇;(ii) 展现出优越泛化性,在有限标注的多标签分类任务中超越7种放射科专用自监督方法;(iii) 融合细粒度信息,仅使用图像级标注即可实现精确异常检测。
原文摘要 · Abstract (English)
Current self-supervised methods, such as contrastive learning, predominantly focus on global discrimination, neglecting the critical fine-grained anatomical details required for accurate radiographic analysis. To address this challenge, we propose an Anatomy-driven self-supervised framework for enhancing Fine-grained Representation in radiographic image analysis (AFiRe). The core idea of AFiRe is to align the anatomical consistency with the unique token-processing characteristics of Vision Transformer. Specifically, AFiRe synergistically performs two self-supervised schemes: (i) Token-wise anatomy-guided contrastive learning, which aligns image tokens based on structural and categorical consistency, thereby enhancing fine-grained spatial-anatomical discrimination; (ii) Pixel-level anomaly-removal restoration, which particularly focuses on local anomalies, thereby refining the learned discrimination with detailed geometrical information. Additionally, we propose Synthetic Lesion Mask to enhance anatomical diversity while preserving intra-consistency, which is typically corrupted by traditional data augmentations, such as Cropping and Affine transformations. Experimental results show that AFiRe: (i) provides robust anatomical discrimination, achieving more cohesive feature clusters compared to state-of-the-art contrastive learning methods; (ii) demonstrates superior generalization, surpassing 7 radiography-specific self-supervised methods in multi-label classification tasks with limited labeling; and (iii) integrates fine-grained information, enabling precise anomaly detection using only image-level annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。