让AI像人一样主动换视角看东西,提升问答准确率
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- 基于当前画面选最有效视角,不依赖记忆或外部知识
- 在合成和真实场景中均显著提升问答准确率
- 可无缝接入现有视觉问答系统,增强探索能力
视觉语言模型(VLMs)在静态图像的视觉问答(VQA)上表现优异,但局限于快照式视觉。相比之下,具身智能体需要主动移动获取更丰富信息。本文提出视觉引导的主动视角选择(VG-AVS)任务:仅根据当前图像中的视觉信息,选择下一个最具信息量的视角,无需依赖场景记忆或外部知识。为此,构建了一个合成数据集,包含自动生成的成对查询-目标视角及问答提示。提出一个框架,通过监督微调(SFT)后接基于强化学习的策略优化,微调预训练的VLM。所提方法在视角选择基础上实现了出色的问答性能,并在未见过的合成与真实场景中表现出强泛化能力。此外,将该框架集成到现有的基于场景探索的EQA系统中,可进一步提升下游问答准确率。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) excel at visual question answering (VQA) but remain limited to snapshot vision, reasoning from static images. In contrast, embodied agents require ambulatory vision, actively moving to obtain more informative views. We introduce Visually Grounded Active View Selection (VG-AVS), a task that selects the most informative next viewpoint using only the visual information in the current image, without relying on scene memory or external knowledge. To support this task, we construct a synthetic dataset with automatically generated paired query-target views and question-answer prompts. We also propose a framework that fine-tunes pretrained VLMs through supervised fine-tuning (SFT) followed by RL-based policy optimization. Our approach achieves strong question answering performance based on viewpoint selection and generalizes robustly to unseen synthetic and real scenes. Furthermore, incorporating our learned VG-AVS framework into existing scene-exploration-based EQA systems improves downstream question-answering accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。