arXiv:2503.07456cs.CV2025-03

基于解剖位置的图文检索,提升罕见病病例匹配精度。

Anatomy-Aware Conditional Image-Text Retrieval

  • 结合解剖区域信息进行图文对齐,增强局部特征匹配
  • 在多个数据集上达到当前最优定位与检索性能
  • 适合临床辅助诊断,支持可解释性初步诊断生成

图像-文本检索(ITR)在医疗领域有广泛应用,可帮助医生高效查找相似患者病例,尤其对罕见病诊疗具有重要意义。然而传统ITR系统仅依赖全局图像或文本表征,忽略病例间的局部差异,导致检索效果不佳。本文提出解剖位置条件化图像-文本检索框架(ALC-ITR),给定查询图像及可疑解剖区域,目标是检索出相同部位出现相似病变或症状的患者病例。为此,我们构建了医学相关性-区域-对齐视觉语言模型(RRA-VL),实现语义层面的全局、区域及词级对齐,生成通用性强的多模态表示。同时引入位置条件对比学习,利用跨样本区域级对比性提升检索能力。实验表明,RRA-VL在相位定位任务中表现领先,且在有无位置条件的情况下均实现优异的多模态检索效果。最后,通过预设LLM提示,系统能基于检索结果提供可解释性的初步诊断报告和解释,验证其泛化性与实用性。

原文摘要 · Abstract (English)

Image-Text Retrieval (ITR) finds broad applications in healthcare, aiding clinicians and radiologists by automatically retrieving relevant patient cases in the database given the query image and/or report, for more efficient clinical diagnosis and treatment, especially for rare diseases. However conventional ITR systems typically only rely on global image or text representations for measuring patient image/report similarities, which overlook local distinctiveness across patient cases. This often results in suboptimal retrieval performance. In this paper, we propose an Anatomical Location-Conditioned Image-Text Retrieval (ALC-ITR) framework, which, given a query image and the associated suspicious anatomical region(s), aims to retrieve similar patient cases exhibiting the same disease or symptoms in the same anatomical region. To perform location-conditioned multimodal retrieval, we learn a medical Relevance-Region-Aligned Vision Language (RRA-VL) model with semantic global-level and region-/word-level alignment to produce generalizable, well-aligned multi-modal representations. Additionally, we perform location-conditioned contrastive learning to further utilize cross-pair region-level contrastiveness for improved multi-modal retrieval. We show that our proposed RRA-VL achieves state-of-the-art localization performance in phase-grounding tasks, and satisfying multi-modal retrieval performance with or without location conditioning. Finally, we thoroughly investigate the generalizability and explainability of our proposed ALC-ITR system in providing explanations and preliminary diagnosis reports given retrieved patient cases (conditioned on anatomical regions), with proper off-the-shelf LLM prompts.

医疗AI图文检索解剖定位可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。