将视觉语言模型与机器人动作模型动态协同,提升真实场景下的自主执行能力。
PhysiAgent: An Embodied Agent Framework in Physical World
- 通过实时反馈动态调度视觉语言模型与动作模型协作
- 在真实机器人任务中显著提升解题成功率
- 适合需要自适应决策的智能机器人系统研究者
视觉-语言-动作(VLA)模型虽取得显著进展,但泛化能力有限。现有方法常将通用视觉语言模型(VLM)作为高层规划者,而将VLA仅用作底层动作执行者,形成僵化、低效的串行结构,导致协作不足与语义漂移问题。本文提出面向物理世界的具身智能体框架PhysiAgent,引入监控、记忆、自我反思机制及轻量级工具箱,构建自主支撑体系。该框架能根据VLA的实时执行表现,动态引导VLM组织任务组件,充分挖掘VLA潜力。实验表明,在复杂真实机器人任务中,PhysiAgent显著提升任务完成性能,实现对VLM的有效自调控、工具间连贯协作以及执行过程中的自适应演化。本工作为融合VLM与VLA提供了切实可行且具有前瞻性的解决方案,有效增强了具身智能体在真实环境中的落地能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations. To address this, integrating generalized Vision-Language Models (VLMs) as assistants to VLAs has emerged as a popular solution. However, current approaches often combine these models in rigid, sequential structures: using VLMs primarily for high-level scene understanding and task planning, and VLAs merely as executors of lower-level actions, leading to ineffective collaboration and poor grounding challenges. In this paper, we propose an embodied agent framework, PhysiAgent, tailored to operate effectively in physical environments. By incorporating monitor, memory, self-reflection mechanisms, and lightweight off-the-shelf toolboxes, PhysiAgent offers an autonomous scaffolding framework to prompt VLMs to organize different components based on real-time proficiency feedback from VLAs to maximally exploit VLAs' capabilities. Experimental results demonstrate significant improvements in task-solving performance on complex real-world robotic tasks, showcasing effective self-regulation of VLMs, coherent tool collaboration, and adaptive evolution of the framework during execution. PhysiAgent makes practical and pioneering efforts to integrate VLMs and VLAs, effectively grounding embodied agent frameworks in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。