让大模型主动提问,提升机器人对环境的理解能力。
PRISM: Perception Reasoning Interleaved for Sequential Decision Making

- 大模型与视觉模型动态交互,主动提问获取关键信息。
- 在ALFWorld和R2R任务中表现超越现有最优模型。
- 无需人工设计问题,全自动化流程适合真实场景应用。
将基于大语言模型的具身智能体从纯文本环境扩展到复杂多模态场景仍是重大挑战。现有研究指出,独立的视觉-语言模型存在感知-推理-决策鸿沟,常忽略任务关键信息。本文提出PRISM框架,通过动态问答管道将视觉模型(VLM)与决策模型(LLM)紧密耦合。不同于被动接受视觉模型的描述,大语言模型会批判性审视、提出目标导向的问题,并合成精炼的图像描述。这种闭环交互实现了精准的任务驱动场景理解。我们在ALFWorld和房间间导航(Room-to-Room, R2R)基准上评估,结果表明:(1) PRISM显著优于当前最先进的基于图像的模型;(2) 我们的交互式目标导向感知管道带来系统性且显著的性能提升;(3) PRISM完全自动,无需人工设计问题或答案。
原文摘要 · Abstract (English)
Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception-reasoning-decision gap in standalone Vision-Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce PRISM, a framework that tightly couples perception (VLM) and decision (LLM) through a dynamic question-answer (DQA) pipeline. Instead of passively accepting the VLM's description, the LLM critiques it, probes the VLM with goal-oriented questions, and synthesizes a compact image description. This closed-loop interaction yields a sharp, task-driven understanding of the scene. We evaluate PRISM on the ALFWorld and Room-to-Room (R2R) benchmarks. We show that: (1) PRISM significantly outperforms state-of-the-art image-based models, (2) our Interactive goal-oriented perception pipeline yields systematic and substantial gains, and (3) PRISM is fully automatic, eliminating the need for handcrafted questions or answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。