构建多模态智能体,实现移动操作的开放世界泛化与自主决策。
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
- 基于多视图场景与状态追踪的跨模态智能体架构,通过函数调用控制机器人
- 在真实世界中达到当前最优性能,且具备强零样本泛化能力
- 首个集成全局场景理解、状态跟踪与多模态动作生成的专用基础模型
导航、操作和视觉模型的快速发展使移动操作机器人能够完成诸多专项任务。然而,开放世界移动操作(OWMM)仍面临挑战,需应对开放式指令与环境的泛化需求,以及将高层决策与低层控制结合的系统复杂性,依赖全局场景理解与当前代理状态。为此,我们提出一种新型多模态智能体架构,通过维护多视角场景帧与代理状态进行决策,并以函数调用方式控制机器人。另一挑战是领域偏移导致的幻觉问题。为提升性能,我们进一步引入用于OWMM任务的智能体数据合成流水线,通过指令微调使视觉语言模型(VLM)适应目标任务域。我们强调,所微调的OWMM-VLM是首个统一集成全局场景理解、机器人状态追踪与多模态动作生成的专用基础模型。实验表明,该模型在真实世界中优于其他基础模型(包括GPT-4o),并展现出强大的零样本泛化能力。
原文摘要 · Abstract (English)
The rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks. However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level decision making with low-level robot control based on both global scene understanding and current agent state. To address this complexity, we propose a novel multi-modal agent architecture that maintains multi-view scene frames and agent states for decision-making and controls the robot by function calling. A second challenge is the hallucination from domain shift. To enhance the agent performance, we further introduce an agentic data synthesis pipeline for the OWMM task to adapt the VLM model to our task domain with instruction fine-tuning. We highlight our fine-tuned OWMM-VLM as the first dedicated foundation model for mobile manipulators with global scene understanding, robot state tracking, and multi-modal action generation in a unified model. Through experiments, we demonstrate that our model achieves SOTA performance compared to other foundation models including GPT-4o and strong zero-shot generalization in real world. The project page is at https://github.com/HHYHRHY/OWMM-Agent
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。