arXiv:2501.09134cs.CVcs.AI2025-01中稿 · AAAI被引 3

对比学习模型在医学影像报告检索中抗干扰能力弱,需改进鲁棒性。

Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval

  • 引入遮挡检索任务评估四种对比学习模型在图像损坏下的表现。
  • 所有模型性能随遮挡程度增加而下降,MedCLIP最稳定但整体仍落后于CXR-CLIP。
  • 通用预训练模型(如CLIP)在医疗领域表现差,需领域专用数据训练。

医学影像与报告为患者健康提供重要信息,但其异质性和复杂性阻碍了有效分析。为弥合这一差距,本文研究用于跨域检索的对比学习模型,将医学影像与其对应临床报告关联。本研究对四种前沿对比学习模型(CLIP、CXR-RePaiR、MedCLIP、CXR-CLIP)进行鲁棒性基准测试。引入遮挡检索任务,评估模型在不同图像损坏水平下的表现。结果表明,所有模型对分布外数据高度敏感,性能随遮挡程度增加呈比例下降。尽管MedCLIP表现出稍强鲁棒性,但其整体性能仍显著低于CXR-CLIP和CXR-RePaiR。CLIP因在通用数据集上训练,在医学图像报告检索中表现不佳,凸显领域专用训练数据的重要性。研究建议需进一步提升模型鲁棒性,以推动更可靠的医疗跨域检索系统发展。

原文摘要 · Abstract (English)

Medical images and reports offer invaluable insights into patient health. The heterogeneity and complexity of these data hinder effective analysis. To bridge this gap, we investigate contrastive learning models for cross-domain retrieval, which associates medical images with their corresponding clinical reports. This study benchmarks the robustness of four state-of-the-art contrastive learning models: CLIP, CXR-RePaiR, MedCLIP, and CXR-CLIP. We introduce an occlusion retrieval task to evaluate model performance under varying levels of image corruption. Our findings reveal that all evaluated models are highly sensitive to out-of-distribution data, as evidenced by the proportional decrease in performance with increasing occlusion levels. While MedCLIP exhibits slightly more robustness, its overall performance remains significantly behind CXR-CLIP and CXR-RePaiR. CLIP, trained on a general-purpose dataset, struggles with medical image-report retrieval, highlighting the importance of domain-specific training data. The evaluation of this work suggests that more effort needs to be spent on improving the robustness of these models. By addressing these limitations, we can develop more reliable cross-domain retrieval models for medical applications.

医学图像对比学习鲁棒性跨模态检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。