通过扰动报告文本提升医学多模态模型语义理解能力
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination

- 用语义结构破坏的文本扰动方法增强跨模态对比学习
- 在多个下游任务中超越强基线,提升表示鲁棒性
- 适合医学图像与报告联合建模的研究者使用
大规模未标注医学图像及其关联报告预训练的视觉-语言模型可学习通用语义表征,助力多种生物医学下游任务。对比学习广泛用于自然图像与标题的预训练,但常见方法常忽视医学文本复杂的领域特异性语义。为此,本文提出一种新方法——扰动报告判别:首先构建一组保持原词但破坏句义结构的文本扰动方法;接着对报告施加不同扰动,让模型基于关联图像判断原始报告与扰动版本的区别。同时,通过对比图像子区域和文本子词的注意力加权特征,提升模型对更高粒度信息的敏感度。在多个下游任务上的实验表明,该方法显著优于现有强基线,学习到更具语义意义且更鲁棒的多模态表征。
原文摘要 · Abstract (English)
Vision-language models pre-trained on large scale of unlabeled biomedical images and associated reports learn generalizable semantic representations. These multi-modal representations can benefit various downstream tasks in the biomedical domain. Contrastive learning is widely used to pre-train vision-language models for general natural images and associated captions. Despite its popularity, we found biomedical texts have complex and domain-specific semantics that are often neglected by common contrastive methods. To address this issue, we propose a novel method, perturbed report discrimination, for pre-train biomedical vision-language models. First, we curate a set of text perturbation methods that keep the same words, but disrupt the semantic structure of the sentence. Next, we apply different types of perturbation to reports, and use the model to distinguish the original report from the perturbed ones given the associated image. Parallel to this, we enhance the sensitivity of our method to higher level of granularity for both modalities by contrasting attention-weighted image sub-regions and sub-words in the image-text pairs. We conduct extensive experiments on multiple downstream tasks, and our method outperforms strong baseline methods. The results demonstrate that our approach learns more semantic meaningful and robust multi-modal representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。