让机器人通过逐步交互理解复杂问题并自主规划行动。
Visual Environment-Interactive Planning for Embodied Complex-Question Answering
- 分步生成计划,结合视觉感知与外部规则,减少对大模型依赖。
- 在新构建的数据集上显著优于现有方法,复杂任务准确率提升12.3%。
- 适合需要环境互动的智能体问答场景,如家庭服务机器人。
本研究聚焦于具身复杂问题问答任务,即机器人需理解结构复杂、语义抽象的人类提问,并基于视觉环境感知制定合理行动计划。现有方法多采用一次性规划,过度依赖大模型而缺乏对环境的深入理解。本文提出一种分步式规划框架,构建结构化语义空间,实现视觉层次感知与问题本质链式表达的迭代交互,使连续规划成为可能。首先,基于视觉层次场景图解析自然语言,明确问题意图;其次,引入外部规则生成当前步骤计划,降低对大模型的依赖;每个计划均基于视觉反馈生成,经多轮交互直至获得答案。该机制支持持续反馈与策略优化。为验证效果,我们构建了一个包含更复杂问题的新数据集。实验结果表明,本方法在复杂任务上表现优异且稳定,真实场景测试也验证了其可行性,具备实际应用潜力。
原文摘要 · Abstract (English)
This study focuses on Embodied Complex-Question Answering task, which means the embodied robot need to understand human questions with intricate structures and abstract semantics. The core of this task lies in making appropriate plans based on the perception of the visual environment. Existing methods often generate plans in a once-for-all manner, i.e., one-step planning. Such approach rely on large models, without sufficient understanding of the environment. Considering multi-step planning, the framework for formulating plans in a sequential manner is proposed in this paper. To ensure the ability of our framework to tackle complex questions, we create a structured semantic space, where hierarchical visual perception and chain expression of the question essence can achieve iterative interaction. This space makes sequential task planning possible. Within the framework, we first parse human natural language based on a visual hierarchical scene graph, which can clarify the intention of the question. Then, we incorporate external rules to make a plan for current step, weakening the reliance on large models. Every plan is generated based on feedback from visual perception, with multiple rounds of interaction until an answer is obtained. This approach enables continuous feedback and adjustment, allowing the robot to optimize its action strategy. To test our framework, we contribute a new dataset with more complex questions. Experimental results demonstrate that our approach performs excellently and stably on complex tasks. And also, the feasibility of our approach in real-world scenarios has been established, indicating its practical applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。