让视觉语言模型更好理解物体视角下的空间关系
Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language Models
- 将物体中心视角的空间推理转化为符号化布局形式
- 在多视角和视觉错觉下表现更鲁棒,提升推理准确率
- 适合需要精准空间关系理解的智能系统开发
视角感知的空间推理涉及从特定视角(如第一人称或物体中心)理解空间关系。尽管视觉语言模型在第一人称场景中表现良好,但在物体中心视角下性能显著下降,因需从物体自身视角推断空间关系。本文提出符号投影布局(SymPL)框架,将物体中心推理重构为视觉语言模型擅长的符号化布局形式。通过投影、抽象、二分与定位四个关键机制,将物体中心问题转化为结构化符号表示。大量实验表明,该重构方法显著提升物体中心与第一人称任务的性能,增强对视觉错觉和多视角场景的鲁棒性,且各组件均贡献关键作用。结果表明,SymPL为复杂视角感知空间推理提供了有效而原则性的解决方案。
原文摘要 · Abstract (English)
Perspective-aware spatial reasoning involves understanding spatial relationships from specific viewpoints-either egocentric (observer-centered) or allocentric (object-centered). While vision-language models (VLMs) perform well in egocentric settings, their performance deteriorates when reasoning from allocentric viewpoints, where spatial relations must be inferred from the perspective of objects within the scene. In this study, we address this underexplored challenge by introducing Symbolic Projective Layout (SymPL), a framework that reformulates allocentric reasoning into symbolic-layout forms that VLMs inherently handle well. By leveraging four key factors-projection, abstraction, bipartition, and localization-SymPL converts allocentric questions into structured symbolic-layout representations. Extensive experiments demonstrate that this reformulation substantially improves performance in both allocentric and egocentric tasks, enhances robustness under visual illusions and multi-view scenarios, and that each component contributes critically to these gains. These results show that SymPL provides an effective and principled approach for addressing complex perspective-aware spatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。