让视觉语言模型学会换位思考,提升对场景的多视角理解能力。
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- 通过抽象场景表征实现视角转换,模拟人类心理图像能力。
- 在合成与真实图像数据集上显著优于现有模型,尤其在跨视角推理任务中。
- 适合需要多智能体协作或复杂环境交互的研究者使用。
我们提出一种基于心理图像模拟的视角感知推理框架,用于提升视觉语言模型(VLMs)的视角理解能力。视角转换是人类高水平视觉理解的关键,对环境互动和与自主代理协作至关重要。尽管当前VLM在空间推理方面取得进展,但普遍缺乏视角感知能力,且严重偏向第一人称视角。为弥合这一差距,我们借鉴人类通过抽象表征进行视角切换的心理机制,提出名为抽象视角变换(APC)的框架。该框架利用目标检测、分割和朝向估计等视觉基础模型构建场景抽象,支持视角转换。在合成与真实图像基准上的实验表明,本方法在视角感知推理任务中显著优于多种VLM及微调的空间推理模型,也超越基于新视图生成的方法。
原文摘要 · Abstract (English)
We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for environmental interaction and collaboration with autonomous agents. Despite advancements in spatial reasoning within VLMs, recent research has shown that modern VLMs significantly lack perspective-aware reasoning capabilities and exhibit a strong bias toward egocentric interpretations. To bridge the gap between VLMs and human perception, we focus on the role of mental imagery, where humans perceive the world through abstracted representations that facilitate perspective shifts. Motivated by this, we propose a framework for perspective-aware reasoning, named Abstract Perspective Change (APC), that effectively leverages vision foundation models, such as object detection, segmentation, and orientation estimation, to construct scene abstractions and enable perspective transformations. Our experiments on synthetic and real-image benchmarks, compared with various VLMs, demonstrate significant improvements in perspective-aware reasoning with our framework, further outperforming fine-tuned spatial reasoning models and novel-view-synthesis-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。