arXiv:2509.22404cs.CVcs.AI2025-09被引 3

用参考图引导视觉语言模型,精准定位医学图像中的器官结构。

RAU: Reference-based Anatomical Understanding with Vision Language Models

  • 通过参考图与目标图的相对空间推理,实现解剖结构识别。
  • 在多个数据集上优于微调SAM2基线,小血管等结构分割更准。
  • 无需大量标注数据,适合临床场景中泛化能力要求高的应用。

基于深度学习的解剖理解对自动报告生成、术中导航和器官定位至关重要,但受限于专家标注数据稀缺。一种有前景的方法是利用已标注参考图像来指导未标注目标图像的解读。尽管近期视觉语言模型(VLMs)展现出一定视觉推理能力,其基于参考的解剖理解与细粒度定位仍有限。本文提出RAU框架,利用VLM实现参考图像引导的解剖理解。我们发现,仅在中等规模数据集上训练的VLM即可通过参考图与目标图间的相对空间关系,识别解剖区域,并在视觉问答(VQA)和边界框预测任务中验证该能力。进一步地,将VLM提取的空间线索与SAM2的细粒度分割能力结合,可实现如血管段等小结构的定位与像素级分割。在两个分布内和两个分布外数据集上,RAU均显著优于使用相同记忆设置的SAM2微调基线,分割更准确,定位更可靠。更重要的是,其强泛化能力使其可扩展至分布外数据,这对医学图像应用极为关键。据我们所知,RAU是首个探索VLM在医学图像中基于参考的解剖结构识别、定位与分割能力的工作,其表现凸显了基于VLM方法在自动化临床流程中的潜力。

原文摘要 · Abstract (English)

Anatomical understanding through deep learning is critical for automatic report generation, intra-operative navigation, and organ localization in medical imaging; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based understanding and fine-grained localization remain limited. We introduce RAU, a framework for reference-based anatomical understanding with VLMs. We first show that a VLM learns to identify anatomical regions through relative spatial reasoning between reference and target images, trained on a moderately sized dataset. We validate this capability through visual question answering (VQA) and bounding box prediction. Next, we demonstrate that the VLM-derived spatial cues can be seamlessly integrated with the fine-grained segmentation capability of SAM2, enabling localization and pixel-level segmentation of small anatomical regions, such as vessel segments. Across two in-distribution and two out-of-distribution datasets, RAU consistently outperforms a SAM2 fine-tuning baseline using the same memory setup, yielding more accurate segmentations and more reliable localization. More importantly, its strong generalization ability makes it scalable to out-of-distribution datasets, a property crucial for medical image applications. To the best of our knowledge, RAU is the first to explore the capability of VLMs for reference-based identification, localization, and segmentation of anatomical structures in medical images. Its promising performance highlights the potential of VLM-driven approaches for anatomical understanding in automated clinical workflows.

解剖理解视觉语言模型医学图像参考图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。