arXiv:2507.12739cs.CVcs.AI2025-07综述

综述基于Transformer的空间定位方法,梳理主流模型与评估标准。

Transformer-based Spatial Grounding: A Comprehensive Survey

  • 系统分析2018-2025年基于Transformer的空间定位方法
  • 总结主流模型架构、常用数据集与评估指标
  • 适合研究多模态对齐与工业落地的开发者参考

空间定位是指将自然语言表达与图像中的相应区域关联起来的任务。得益于Transformer模型的引入,该领域在多模态表征和跨模态对齐方面取得显著进展。尽管如此,当前领域仍缺乏对现有方法、数据集使用、评估指标及工业适用性的全面综述。本文对2018至2025年间基于Transformer的空间定位方法进行了系统性文献回顾。分析揭示了主导的模型架构、流行的训练数据集以及广泛采用的评价指标,并指出了关键的方法趋势与最佳实践。本研究为研究人员与从业者提供了重要洞见与结构化指导,有助于开发稳健、可靠且可工业部署的Transformer基空间定位模型。

原文摘要 · Abstract (English)

Spatial grounding, the process of associating natural language expressions with corresponding image regions, has rapidly advanced due to the introduction of transformer-based models, significantly enhancing multimodal representation and cross-modal alignment. Despite this progress, the field lacks a comprehensive synthesis of current methodologies, dataset usage, evaluation metrics, and industrial applicability. This paper presents a systematic literature review of transformer-based spatial grounding approaches from 2018 to 2025. Our analysis identifies dominant model architectures, prevalent datasets, and widely adopted evaluation metrics, alongside highlighting key methodological trends and best practices. This study provides essential insights and structured guidance for researchers and practitioners, facilitating the development of robust, reliable, and industry-ready transformer-based spatial grounding models.

空间定位Transformer多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。