利用影像多视角结构自监督训练,提升医学影像表征能力
Structure is Supervision: Multiview Masked Autoencoders for Radiology
- 通过掩码重建与跨视角对齐,挖掘影像多视图内在结构作为监督信号
- 在三个公开数据集上优于有监督和视觉-语言模型,低标签场景下表现更优
- 适合构建可泛化、临床可解释的医学基础模型,尤其适用于标注稀缺场景
构建鲁棒的医学机器学习系统需要利用临床数据中固有的结构进行预训练。我们提出多视角掩码自编码器(MVMAE),一种自监督框架,利用放射学检查中天然的多视图组织结构,学习视角不变且与疾病相关的表征。MVMAE结合掩码图像重建与跨视图对齐,将不同投影间的临床冗余转化为强大的自监督信号。我们进一步提出MVMAE-V2T,引入放射科报告作为辅助文本信号,在保持全视觉推理的同时增强语义对齐。在三个大规模公开数据集MIMIC-CXR、CheXpert和PadChest上的下游疾病分类任务中,MVMAE consistently 超过有监督及视觉-语言基线模型。此外,MVMAE-V2T在低标签场景下带来额外提升,表明结构化与文本监督是构建可扩展、临床可信医学基础模型的互补路径。
原文摘要 · Abstract (English)
Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages the natural multi-view organization of radiology studies to learn view-invariant and disease-relevant representations. MVMAE combines masked image reconstruction with cross-view alignment, transforming clinical redundancy across projections into a powerful self-supervisory signal. We further extend this approach with MVMAE-V2T, which incorporates radiology reports as an auxiliary text-based learning signal to enhance semantic grounding while preserving fully vision-based inference. Evaluated on a downstream disease classification task on three large-scale public datasets, MIMIC-CXR, CheXpert, and PadChest, MVMAE consistently outperforms supervised and vision-language baselines. Furthermore, MVMAE-V2T provides additional gains, particularly in low-label regimes where structured textual supervision is most beneficial. Together, these results establish the importance of structural and textual supervision as complementary paths toward scalable, clinically grounded medical foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。