arXiv:2603.26646cs.CV2026-03被引 1

用手指指物的视觉定位,让机器更懂人类真实交互。

Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision

  • 构建首个基于第一视角的手指指物数据集,含1.5万组多粒度标注。
  • 新方法在复杂场景中提升定位准确率11.7%,显著减少语义歧义。
  • 适合研究人机交互、多模态理解与具身智能的开发者参考。

传统视觉定位主要依赖文本描述,难以应对语言模糊性,且忽略真实交互中常见的非语言指示行为。在自然的第一视角互动中,手势与语音结合是最直观的指代方式。为此,我们提出 EgoPoint-Ground,首个面向第一视角指物行为的大规模多模态数据集,包含超过15,000个复杂场景中的交互样本,提供手部目标边界框对与密集语义描述等多粒度标注。我们建立了一个全面的基准,评估主流多模态大模型(MLLMs)和先进视觉定位架构。同时提出SV-CoT新框架,将定位任务重构为结构化推理过程,通过视觉思维链融合手势与语言线索。大量实验表明,该方法相比现有方法提升11.7%的绝对性能,有效缓解语义歧义,增强智能体理解多模态物理意图的能力。数据集与代码将公开发布。

原文摘要 · Abstract (English)

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in real-world interactions. In natural egocentric engagements, hand-pointing combined with speech forms the most intuitive referring mechanism. To bridge this gap, we introduce EgoPoint-Ground, the first large-scale multimodal dataset dedicated to egocentric deictic visual grounding. Comprising over \textbf{15k} interactive samples in complex scenes, the dataset provides rich, multi-grained annotations including hand-target bounding box pairs and dense semantic captions. We establish a comprehensive benchmark for hand-pointing referring expression resolution, evaluating a wide spectrum of mainstream Multimodal Large Language Models (MLLMs) and state-of-the-art VG architectures. Furthermore, we propose SV-CoT, a novel baseline framework that reformulates grounding as a structured inference process, synergizing gestural and linguistic cues through a Visual Chain-of-Thought paradigm. Extensive experiments demonstrate that SV-CoT achieves an $\textbf{11.7\%}$ absolute improvement over existing methods, effectively mitigating semantic ambiguity and advancing the capability of agents to comprehend multimodal physical intents. The dataset and code will be made publicly available.

视觉定位多模态手势识别具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。