arXiv:2603.19451cs.CVcs.AI2026-03

通过位置感知机制提升胸部X光片细粒度表征,增强病灶定位与检索能力。

LoFi: Location-Aware Fine-Grained Representation Learning for Chest X-ray

  • 引入位置感知字幕损失,实现区域级监督。
  • 在MIMIC-CXR和PadChest-GR上达到最优检索与定位性能。
  • 适合需要精准病灶定位的临床辅助诊断场景。

细粒度表征学习对胸部X光片的检索与短语定位至关重要,因临床发现常具空间局限性。然而,对比模型缺乏区域级监督,大型视觉语言模型在外部验证中捕捉细粒度表征的能力有限,导致性能不佳。为此,我们提出位置感知细粒度表征学习(LoFi),利用轻量级大语言模型联合优化Sigmoid、字幕生成与位置感知字幕损失。位置感知字幕损失通过定位与密集字幕目标实现区域级监督,促进细粒度表征学习。在此基础上,我们将细粒度编码器集成至基于检索的上下文学习中,提升不同场景下胸部X光片的定位能力。大量实验表明,该方法在MIMIC-CXR和PadChest-GR数据集上均取得优异的检索与短语定位表现。

原文摘要 · Abstract (English)

Fine-grained representation learning is crucial for retrieval and phrase grounding in chest X-rays, where clinically relevant findings are often spatially confined. However, the lack of region-level supervision in contrastive models and the limited ability of large vision language models to capture fine-grained representations in external validation lead to suboptimal performance on these tasks. To address these limitations, we propose Location-aware Fine-grained representation learning (LoFi), which jointly optimizes sigmoid, captioning, and location-aware captioning losses using a lightweight large language model. The location-aware captioning loss enables region-level supervision through grounding and dense captioning objectives, thereby facilitating fine-grained representation learning. Building upon these representations, we integrate a fine-grained encoder into retrieval-based in-context learning to enhance chest X-ray grounding across diverse settings. Extensive experiments demonstrate that our method achieves superior retrieval and phrase grounding performance on MIMIC-CXR and PadChest-GR.

医学影像细粒度学习定位识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。