分层对齐影像与文本,提升医学X光图表示学习效果
Advancing Medical Radiograph Representation Learning: A Hybrid Pre-training Paradigm with Multilevel Semantic Granularity
- 分两层对齐:整体图像与报告结论、局部图像与病灶描述
- 生成式解码器通过摘要和描述任务提升模型性能
- 知识蒸馏使模型更高效,参数增长小但效果显著
本文针对医学影像-文本预训练领域中的放射科图像表示学习问题,提出一种新型混合预训练框架HybridMED。传统方法常将文本注释合并为统一报告,而我们注意到报告中发现与结论之间存在固有层次关系。为此,框架分别对齐全局视觉特征与结论,以及词级视觉特征与病灶描述。同时引入生成解码器,包含两个代理任务:从图像生成结论(描述分支),从病灶生成结论(摘要分支)。此外,采用知识蒸馏策略优化训练过程。在MIMIC-CXR数据集上的实验表明,摘要分支能有效向描述分支传递知识,显著提升模型表现,且因共享自注意力与前馈网络结构,参数增加极少。
原文摘要 · Abstract (English)
This paper introduces an innovative approach to Medical Vision-Language Pre-training (Med-VLP) area in the specialized context of radiograph representation learning. While conventional methods frequently merge textual annotations into unified reports, we acknowledge the intrinsic hierarchical relationship between the findings and impression section in radiograph datasets. To establish a targeted correspondence between images and texts, we propose a novel HybridMED framework to align global-level visual representations with impression and token-level visual representations with findings. Moreover, our framework incorporates a generation decoder that employs two proxy tasks, responsible for generating the impression from (1) images, via a captioning branch, and (2) findings, through a summarization branch. Additionally, knowledge distillation is leveraged to facilitate the training process. Experiments on the MIMIC-CXR dataset reveal that our summarization branch effectively distills knowledge to the captioning branch, enhancing model performance without significantly increasing parameter requirements due to the shared self-attention and feed-forward architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。