arXiv:2502.03118cs.CVcs.AI2025-02

用同一段语言提示让模型自动对齐图像,无需训练即可完成精准配准。

Tell2Reg: Establishing spatial correspondence between images by the same language prompts

  • 用相同语言提示触发双图语义匹配,实现区域级空间对应。
  • 在前列腺多患者MRI配准任务中超越无监督方法,接近弱监督性能。
  • 首次揭示语言语义与空间对应存在关联,适合医学图像跨模态配准。

空间对应可通过成对分割区域表示,图像配准网络目标是分割对应区域而非预测位移场或变换参数。本文表明,利用基于GroundingDINO和SAM的预训练大模型,对两张不同图像使用相同的语言提示,可准确预测对应的区域对。该方法实现了完全自动化且无需训练的配准算法,具备广泛适用性。实验以高挑战性的跨受试者前列腺MRI图像配准为例,该任务存在显著的强度与形态差异。Tell2Reg无需训练,避免了以往所需的昂贵且耗时的数据清洗与标注。其性能超越测试的无监督学习方法,达到与弱监督方法相当的水平。额外定性结果表明,首次发现语言语义与空间对应间存在潜在关联,包括语言提示区域的空间不变性,以及局部与全局对应中语言提示的差异。

原文摘要 · Abstract (English)

Spatial correspondence can be represented by pairs of segmented regions, such that the image registration networks aim to segment corresponding regions rather than predicting displacement fields or transformation parameters. In this work, we show that such a corresponding region pair can be predicted by the same language prompt on two different images using the pre-trained large multimodal models based on GroundingDINO and SAM. This enables a fully automated and training-free registration algorithm, potentially generalisable to a wide range of image registration tasks. In this paper, we present experimental results using one of the challenging tasks, registering inter-subject prostate MR images, which involves both highly variable intensity and morphology between patients. Tell2Reg is training-free, eliminating the need for costly and time-consuming data curation and labelling that was previously required for this registration task. This approach outperforms unsupervised learning-based registration methods tested, and has a performance comparable to weakly-supervised methods. Additional qualitative results are also presented to suggest that, for the first time, there is a potential correlation between language semantics and spatial correspondence, including the spatial invariance in language-prompted regions and the difference in language prompts between the obtained local and global correspondences. Code is available at https://github.com/yanwenCi/Tell2Reg.git.

图像配准语言引导零样本医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。