arXiv:2608.15605cs.CV2026-08

让视觉语言模型分清空间方向是全局还是自身视角

AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models

论文配图:AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models
图 1 · 摘自论文原文
  • 构建新数据集,标注物体在全局与自身视角下的空间关系
  • 在模糊指令下仍能准确判断方向,提升机器人导航能力
  • 可无缝接入现有模型,适合做具身智能的开发者

本研究针对视觉语言模型(VLMs)在理解空间语义时面临的歧义问题。空间认知受认知心理学、空间科学和文化背景影响,常赋予物体方向性,但自然语言描述中常省略参考坐标系,导致语义模糊,对具身AI机器人可能造成严重错误。现有VLM因缺乏对参考框架和物体朝向的充分训练,常产生不一致回答。为此,我们构建了新数据集AlloEgo-View,包含(图像、查询、视图特定答案)三元组,从全局和自身视角捕捉关键物体关系。该数据集采用结构化空间表示,标注详细场景描述、参考与目标物体、其朝向、参考框架和视图类型。基于此,我们提出AlloEgo-VLM框架,可在模糊查询下准确区分全局与自身参考框架,并可通过监督微调轻松集成至现有VLM。进一步在NVIDIA Isaac Sim中部署于具身机器人平台,验证其在开放域物体搜索任务中的实际可行性。实验表明当前VLM在视图相关查询上存在明显局限,而AlloEgo-VLM展现出强歧义消解能力。

原文摘要 · Abstract (English)

This study investigates the challenge of ambiguity faced by Vision-Language Models (VLMs) in understanding spatial semantics. Spatial cognition, shaped by cognitive psychology, spatial science, and cultural context, often assigns directionality to objects. However, natural language descriptions of spatial relations frequently omit explicit reference frames, leading to semantic ambiguity and potentially serious errors for embodied AI robots. Existing VLMs, due to insufficient training on reference frames and object orientations, often produce inconsistent responses. To address this issue, we construct a new dataset, AlloEgo-View, comprising (image, query, view-specific answer) triplets that capture key object relations from both allocentric and egocentric perspectives. The view-specific descriptions follow a structured spatial representation that annotate detailed scene descriptions, reference and target objects, their orientations, reference frames, and view types. Building on AlloEgo-View, we develop AlloEgo-VLM, a framework to disambiguate allocentric and egocentric reference frames, even under ambiguous queries, and to be easily integrated into existing VLMs via supervised fine-tuning. Furthermore, we deploy our framework onto an embodied robotic platform within NVIDIA Isaac Sim to validate its real-world feasibility in open-ended object searching tasks. Experiments highlight the limitations of current VLMs in handling view-specific queries and demonstrate the strong disambiguation ability of AlloEgo-VLM.

空间理解具身智能视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。