arXiv:2511.08098cs.ROcs.AI2025-11中稿 · IAS19被引 3

让大模型学会换位思考,提升协作机器人表现

PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision

  • 用ReAct框架结合主动视觉探索,训练模型理解他人视角
  • 在7个复杂场景中,准确率显著提升,解决指代模糊问题
  • 适合研究多智能体协作与具身智能的开发者

大型语言模型(LLMs)和多模态基础模型的进展极大拓展了其在机器人和协同系统中的应用。然而,有效多智能体交互需要强大的换位思考能力,使模型能够理解物理与认知层面的视角差异。现有训练范式常忽视交互上下文,导致模型在推理个体主观视角或处理多观察者环境时面临挑战。本研究评估了在ReAct框架(融合推理与行动)中显式引入多元视角是否能增强LLM理解并回应其他智能体需求的能力。我们通过扩展经典导演任务,在七种逐步增加换位思考复杂度的场景中引入主动视觉探索。这些场景考验智能体基于视觉可及性与交互,解决指代模糊问题的能力,涉及不同状态表示与提示策略,包括ReAct式推理。结果表明,显式视角线索与主动探索策略的结合,显著提升了模型的解释准确性与协作效率。研究揭示了将主动感知与换位思考机制结合,在推动LLM于机器人与多智能体系统中应用方面的潜力,为未来自适应、情境感知人工智能系统的研究奠定基础。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) and multimodal foundation models have significantly broadened their application in robotics and collaborative systems. However, effective multi-agent interaction necessitates robust perspective-taking capabilities, enabling models to interpret both physical and epistemic viewpoints. Current training paradigms often neglect these interactive contexts, resulting in challenges when models must reason about the subjectivity of individual perspectives or navigate environments with multiple observers. This study evaluates whether explicitly incorporating diverse points of view using the ReAct framework, an approach that integrates reasoning and acting, can enhance an LLM's ability to understand and ground the demands of other agents. We extend the classic Director task by introducing active visual exploration across a suite of seven scenarios of increasing perspective-taking complexity. These scenarios are designed to challenge the agent's capacity to resolve referential ambiguity based on visual access and interaction, under varying state representations and prompting strategies, including ReAct-style reasoning. Our results demonstrate that explicit perspective cues, combined with active exploration strategies, significantly improve the model's interpretative accuracy and collaborative effectiveness. These findings highlight the potential of integrating active perception with perspective-taking mechanisms in advancing LLMs' application in robotics and multi-agent systems, setting a foundation for future research into adaptive and context-aware AI systems.

多智能体换位思考具身智能视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。