arXiv:2602.17689cs.LGcs.AI2026-02

让医疗视觉语言模型更抗设备差异,提升真实场景下的可靠性。

Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction

  • 通过自监督学习显式建模鲁棒性,融合扰动感知掩码与跨域一致性约束。
  • 在多个医学多模态任务中表现领先,跨域VQA准确率最高达78.9%。
  • 特别适合部署于不同设备、协议的临床实际场景,提升诊断可信度。

医疗视觉语言模型在联合推理医学图像与临床文本方面潜力巨大,但受成像设备、采集协议和报告风格差异影响,性能常因领域偏移而下降。现有方法多忽略鲁棒性,将其视为下游适应问题。本文提出鲁棒多模态掩码重建(Robust-MMR),一种自监督预训练框架,将鲁棒性目标显式融入掩码视觉语言学习中。Robust-MMR结合非对称扰动感知掩码、域一致性正则化与模态韧性约束,促进域不变表征。在多个医学多模态基准上评估,包括医学视觉问答(VQA-RAD、SLAKE、VQA-2019)、跨域图文分类(MELINDA)及鲁棒图文检索(ROCO)。Robust-MMR在VQA-RAD上达到78.9%跨域准确率,比最强基线高3.8个百分点;在SLAKE和VQA-2019上分别达74.6%和77.0%。在扰动测试下,其VQA-RAD准确率从69.1%提升至75.6%。图文分类中,跨域MELINDA准确率由70.3%升至75.2%;检索实验显示,平均排名退化从超16降至4.1。定性分析进一步证明其在疾病检测与结构异常评估中具备更强临床推理能力。结果表明,预训练阶段显式建模鲁棒性可获得更可靠、可迁移的医疗视觉语言表征,适用于真实世界部署。

原文摘要 · Abstract (English)

Medical vision-language models show strong potential for joint reasoning over medical images and clinical text, but their performance often degrades under domain shift caused by variations in imaging devices, acquisition protocols, and reporting styles. Existing multi-modal pre-training methods largely overlook robustness, treating it as a downstream adaptation problem. In this work, we propose Robust Multi-Modal Masked Reconstruction (Robust-MMR), a self-supervised pre-training framework that explicitly incorporates robustness objectives into masked vision-language learning. Robust-MMR integrates asymmetric perturbation-aware masking, domain-consistency regularization, and modality-resilience constraints to encourage domain-invariant representations. We evaluate Robust-MMR on multiple medical vision-language benchmarks, including medical visual question answering (VQA-RAD, SLAKE, VQA-2019), cross-domain image-text classification (MELINDA), and robust image-caption retrieval (ROCO). Robust-MMR achieves 78.9% cross-domain accuracy on VQA-RAD, outperforming the strongest baseline by 3.8 percentage points, and reaches 74.6% and 77.0% accuracy on SLAKE and VQA-2019, respectively. Under perturbed evaluation, Robust-MMR improves VQA-RAD accuracy from 69.1% to 75.6%. For image-text classification, cross-domain MELINDA accuracy increases from 70.3% to 75.2%, while retrieval experiments show a reduction in mean rank degradation from over 16 to 4.1 under perturbation. Qualitative results further demonstrate improved clinical reasoning for disease detection and structural abnormality assessment. These findings show that explicitly modeling robustness during pre-training leads to more reliable and transferable medical vision-language representations for real-world deployment.

医疗多模态鲁棒性自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。