用大模型规划+3D技能策略,让机器人在厨房更智能地完成复杂操作。
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
- 大模型动态理解环境,失败后可重试并记忆历史策略。
- 3D语义特征场提升操控精度,低层控制成功率提高45%。
- 适合需要长时序、多步骤的现实场景机器人任务研究者。
大型多模态模型(LMM)在视觉推理能力上的进展,以及3D特征场的语义增强,极大地拓展了机器人的能力边界。本文提出LMM-3DP框架,实现LMM规划器与3D技能策略的有效融合。该方法包含三个关键部分:高层规划、底层控制与高效集成。高层规划支持动态环境感知、带自反馈的评判代理、历史策略记忆及失败后的重试机制。底层控制采用语义感知的3D特征场,实现精准操作。通过在3D变换器中联合关注语言嵌入与3D特征场,实现高/低层控制的无缝衔接。我们在真实厨房环境中对多种技能和长时序任务进行了全面评估。结果表明,相比基于LLM的基线方法,低层控制成功率提升1.45倍,高层规划准确率约提升1.5倍。演示视频与框架概览见https://lmm-3dp-release.github.io。
原文摘要 · Abstract (English)
The recent advancements in visual reasoning capabilities of large multimodal models (LMMs) and the semantic enrichment of 3D feature fields have expanded the horizons of robotic capabilities. These developments hold significant potential for bridging the gap between high-level reasoning from LMMs and low-level control policies utilizing 3D feature fields. In this work, we introduce LMM-3DP, a framework that can integrate LMM planners and 3D skill Policies. Our approach consists of three key perspectives: high-level planning, low-level control, and effective integration. For high-level planning, LMM-3DP supports dynamic scene understanding for environment disturbances, a critic agent with self-feedback, history policy memorization, and reattempts after failures. For low-level control, LMM-3DP utilizes a semantic-aware 3D feature field for accurate manipulation. In aligning high-level and low-level control for robot actions, language embeddings representing the high-level policy are jointly attended with the 3D feature field in the 3D transformer for seamless integration. We extensively evaluate our approach across multiple skills and long-horizon tasks in a real-world kitchen environment. Our results show a significant 1.45x success rate increase in low-level control and an approximate 1.5x improvement in high-level planning accuracy compared to LLM-based baselines. Demo videos and an overview of LMM-3DP are available at https://lmm-3dp-release.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。