让机器人在动态环境中自适应操作,靠世界模型+扩散策略闭环更新。
AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation
- 用世界模型驱动的扩散策略,结合力反馈实现在线自适应。
- 在模拟与真实机器人任务中均达顶尖性能,对分布外场景有强鲁棒性。
- 适合需要高自适应能力的复杂机械臂操作任务研究者。
有效的机器人操作需要能够预判物理结果并适应现实环境的策略。本文提出统一框架AdaWorldPolicy——基于世界模型的扩散策略与在线自适应学习,以在动态条件下实现最小人工干预的机器人操作。核心思想是:世界模型提供强监督信号,支持动态环境中的在线自适应学习,并可结合力-扭矩反馈缓解动态受力变化。该方法集成世界模型、动作专家与力预测器,全部采用互联的流匹配扩散变换器(DiT),通过多模态自注意力层实现深度特征交互,同时保持模块独立性。进一步提出新型在线自适应学习(AdaOL)策略,动态切换动作生成与未来想象模式,推动三个模块协同更新。形成高效闭环机制,可低开销适应视觉与物理域偏移。在一系列模拟与真实机器人基准测试中,本方法达到当前最优表现,具备出色的分布外场景适应能力。
原文摘要 · Abstract (English)
Effective robotic manipulation requires policies that can anticipate physical outcomes and adapt to real-world environments. Effective robotic manipulation requires policies that can anticipate physical outcomes and adapt to real-world environments. In this work, we introduce a unified framework, World-Model-Driven Diffusion Policy with Online Adaptive Learning (AdaWorldPolicy) to enhance robotic manipulation under dynamic conditions with minimal human involvement. Our core insight is that world models provide strong supervision signals, enabling online adaptive learning in dynamic environments, which can be complemented by force-torque feedback to mitigate dynamic force shifts. Our AdaWorldPolicy integrates a world model, an action expert, and a force predictor-all implemented as interconnected Flow Matching Diffusion Transformers (DiT). They are interconnected via the multi-modal self-attention layers, enabling deep feature exchange for joint learning while preserving their distinct modularity characteristics. We further propose a novel Online Adaptive Learning (AdaOL) strategy that dynamically switches between an Action Generation mode and a Future Imagination mode to drive reactive updates across all three modules. This creates a powerful closed-loop mechanism that adapts to both visual and physical domain shifts with minimal overhead. Across a suite of simulated and real-robot benchmarks, our AdaWorldPolicy achieves state-of-the-art performance, with dynamical adaptive capacity to out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。