arXiv:2412.20206cs.CV2024-12TPAMI综述被引 65

全面梳理视觉定位十年进展,涵盖新范式与未来方向。

Toward Visual Grounding: A Survey

  • 系统归纳视觉定位发展脉络与核心设定
  • 覆盖从基础任务到大像素、多模态大模型等前沿方向
  • 适合初学者入门与研究者追踪最新挑战

视觉定位(Visual Grounding),又称指代表达理解与短语定位,旨在根据给定的文本表达,在图像中定位特定区域。该任务模拟了视觉与语言模态间的常见指代关系,使机器具备类人般的多模态理解能力,具有广泛应用前景。自2021年以来,该领域取得显著进展,涌现出基于定位的预训练、多模态大模型定位、泛化视觉定位和千兆像素级定位等新概念,带来诸多新挑战。本文首先回顾视觉定位的发展历程,梳理基础背景知识;系统追踪并总结技术演进,严谨定义并组织各类任务设置,以规范未来研究并确保公平比较;深入分析相关数据集与应用,突出若干前沿主题;最后指出当前面临的关键挑战,并提出有价值的研究方向。通过提炼共性技术细节,本综述涵盖过去十年各子领域的代表性工作。据我们所知,这是目前该领域最全面的综述。本文适用于初学者与资深研究者,是理解核心概念与跟踪最新进展的重要资源。相关工作持续更新于 https://github.com/linhuixiao/Awesome-Visual-Grounding。

原文摘要 · Abstract (English)

Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships between visual and linguistic modalities, enabling machines to develop human-like multimodal comprehension capabilities. Consequently, it has extensive applications in various domains. However, since 2021, visual grounding has witnessed significant advancements, with emerging new concepts such as grounded pre-training, grounding multimodal LLMs, generalized visual grounding, and giga-pixel grounding, which have brought numerous new challenges. In this survey, we first examine the developmental history of visual grounding and provide an overview of essential background knowledge. We systematically track and summarize the advancements, and then meticulously define and organize the various settings to standardize future research and ensure a fair comparison. Additionally, we delve into numerous related datasets and applications, and highlight several advanced topics. Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may serve as inspiration for subsequent researchers. By extracting common technical details, this survey encompasses the representative work in each subtopic over the past decade. To the best of our knowledge, this paper represents the most comprehensive overview currently available in the field of visual grounding. This survey is designed to be suitable for both beginners and experienced researchers, serving as an invaluable resource for understanding key concepts and tracking the latest research developments. We keep tracing related work at https://github.com/linhuixiao/Awesome-Visual-Grounding.

视觉定位多模态综述指代表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。