arXiv:2507.02987cs.CVcs.LG2025-07被引 2

利用胸部X光片的结构关系提升医疗图像表征学习效果

Leveraging the Structure of Medical Data for Improved Representation Learning

  • 用前后位与侧位胸片作为正样本对,自监督学习图像重建与嵌入对齐
  • 在MIMIC-CXR数据集上性能超越有监督基线和未利用结构的模型
  • 适合数据有限但具有成对结构的医疗领域预训练任务

构建可泛化的医疗AI系统需要高效且具备领域感知的预训练策略。与互联网规模语料不同,临床数据集如MIMIC-CXR图像数量有限、标注稀少,但通过多视角成像展现出丰富的内部结构。本文提出一种自监督框架,利用医疗数据的内在结构:将配对的胸片(前后位与侧位)视为自然正样本对,通过稀疏补丁重建每个视图并对其潜在嵌入进行对齐。该方法无需文本监督,生成具信息量的表示。在MIMIC-CXR上的评估表明,其性能显著优于有监督目标及未利用结构的基线模型。本工作为数据结构化但稀缺的领域提供了轻量、模态无关的预训练范式。

原文摘要 · Abstract (English)

Building generalizable medical AI systems requires pretraining strategies that are data-efficient and domain-aware. Unlike internet-scale corpora, clinical datasets such as MIMIC-CXR offer limited image counts and scarce annotations, but exhibit rich internal structure through multi-view imaging. We propose a self-supervised framework that leverages the inherent structure of medical datasets. Specifically, we treat paired chest X-rays (i.e., frontal and lateral views) as natural positive pairs, learning to reconstruct each view from sparse patches while aligning their latent embeddings. Our method requires no textual supervision and produces informative representations. Evaluated on MIMIC-CXR, we show strong performance compared to supervised objectives and baselines being trained without leveraging structure. This work provides a lightweight, modality-agnostic blueprint for domain-specific pretraining where data is structured but scarce

医疗表征自监督学习多视角对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。