梳理视觉语言模型中视觉定位的核心技术与应用前景
Towards Understanding Visual Grounding in Visual Language Models
- 系统梳理视觉定位在多模态模型中的实现机制
- 总结主流评估基准与评价指标体系
- 适合关注多模态理解与推理的研究者阅读
视觉定位指模型识别图像中与文本描述匹配区域的能力,具备该能力的模型可支持指代表达理解、细粒度图像/视频问答、显式指代实体的图文生成,以及模拟与真实环境中的低阶和高阶控制。本文综述现代通用视觉语言模型(VLMs)在视觉定位领域的代表性研究,首先阐明定位的重要性,进而解析当代构建定位模型的核心组件,考察其在实际应用中的表现,包括基准测试与评估指标。同时探讨视觉定位、多模态思维链与推理之间的复杂关联。最后分析当前视觉定位面临的关键挑战,并提出未来研究的可行方向。
原文摘要 · Abstract (English)
Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in various domains, including referring expression comprehension, answering questions pertinent to fine-grained details in images or videos, caption visual context by explicitly referring to entities, as well as low and high-level control in simulated and real environments. In this survey paper, we review representative works across the key areas of research on modern general-purpose vision language models (VLMs). We first outline the importance of grounding in VLMs, then delineate the core components of the contemporary paradigm for developing grounded models, and examine their practical applications, including benchmarks and evaluation metrics for grounded multimodal generation. We also discuss the multifaceted interrelations among visual grounding, multimodal chain-of-thought, and reasoning in VLMs. Finally, we analyse the challenges inherent to visual grounding and suggest promising directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。