arXiv:2608.18309cs.CV2026-08

用视觉语言模型实现不同显微图像的精准定位,解决跨模态匹配难题。

XRF-to-Optical Field-of-View Localization with Vision Language Models

论文配图:XRF-to-Optical Field-of-View Localization with Vision Language Models
图 1 · 摘自论文原文
  • 无需训练,利用视觉语言模型直接推理跨模态位置关系。
  • 在低对应性样本中,通过候选生成与验证流程实现有效定位。
  • 适合需要跨模态配准的生物医学图像分析研究者使用。

不同显微成像模态间的图像配准对于关联同一标本的互补测量至关重要。在共焦X射线荧光(XRF)与光学显微镜成像中,XRF图通常仅覆盖光学图像的一小部分区域,且来自同一或相邻组织切片。由于模态间外观和结构差异,视野(FOV)定位具有挑战性。本文在两个数据集上评估了无需训练的视觉语言模型(VLM)定位方法,分别代表同切片高对应性和邻切片低对应性成像。测试了无约束和元数据约束搜索,并将VLM与几何控制、经典模板匹配及两种替代的无训练方法(DINOv2和multiGradICON)进行了比较。直接提示VLM产生依赖内容的空间信号,但不可靠;经典匹配在跨模态结构保持时最准确,但在低对应性数据集中失败。提出‘提案-验证’工作流,利用重复VLM预测生成候选位置,并通过图像相似性筛选最终结果,在低对应性条件下仍实现有效定位。

原文摘要 · Abstract (English)

Registering images acquired with different microscopy modalities is essential for relating complementary measurements of the same specimen. In correlative X-ray fluorescence (XRF) and optical microscopy, the XRF map often covers only a small region of an optical image acquired from the same or an adjacent tissue section. Field-of-view (FOV) localization is necessary but can be difficult when appearance and structure differ across modalities. Here we evaluate training-free vision language model (VLM) localization on two datasets representing same-section high-correspondence and adjacent-section low-correspondence imaging. We test unconstrained and metadata-constrained search and compare VLMs with geometric controls, classical template matching, and two alternative training-free approaches (DINOv2 and multiGradICON). Direct VLM prompting produced content-dependent spatial signals but was not reliable alone. Classical matching was most accurate when cross-modal structure was preserved but failed in the low-correspondence collection. A proposal-and-verify workflow used repeated VLM predictions as candidates and image-based similarity to select the final location. This workflow recovered useful localization in the low-correspondence regime.

跨模态对齐视觉语言模型显微成像图像配准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。