用原型增强模型提升医学图文检索的准确性和鲁棒性。
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
- 引入多层级原型捕捉医学数据语义多样性。
- 通过双流置信度估计,最高提升10.17%检索精度。
- 适合临床场景下复杂模糊数据的可靠检索应用。
在跨模态检索任务(如图像到报告、报告到图像检索)中,准确对齐医学图像与相关文本报告至关重要,但因医学数据固有的模糊性和变异性而极具挑战。现有模型难以捕捉放射科数据中的细微多层级语义关系,导致检索结果不可靠。为此,我们提出原型增强置信度建模框架(PECM),为每种模态引入多层级原型,更好捕捉语义变异性并增强检索鲁棒性。PECM采用双流置信度估计,利用原型相似度分布和自适应加权机制,控制高不确定性数据对排序的影响。在放射科图像-报告数据集上,该方法显著提升检索精度与一致性,有效应对数据模糊性,在多种数据集和任务(包括全监督与零样本检索)中实现最高10.17%的性能提升,建立新基准。
原文摘要 · Abstract (English)
In cross-modal retrieval tasks, such as image-to-report and report-to-image retrieval, accurately aligning medical images with relevant text reports is essential but challenging due to the inherent ambiguity and variability in medical data. Existing models often struggle to capture the nuanced, multi-level semantic relationships in radiology data, leading to unreliable retrieval results. To address these issues, we propose the Prototype-Enhanced Confidence Modeling (PECM) framework, which introduces multi-level prototypes for each modality to better capture semantic variability and enhance retrieval robustness. PECM employs a dual-stream confidence estimation that leverages prototype similarity distributions and an adaptive weighting mechanism to control the impact of high-uncertainty data on retrieval rankings. Applied to radiology image-report datasets, our method achieves significant improvements in retrieval precision and consistency, effectively handling data ambiguity and advancing reliability in complex clinical scenarios. We report results on multiple different datasets and tasks including fully supervised and zero-shot retrieval obtaining performance gains of up to 10.17%, establishing in new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。