用双塔4D模型实现高保真仿真,统一优化机器人策略。
RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
- 双塔结构+跨模态增强,保证时空几何一致性
- 在精细操作任务上策略提升超97%
- 适合需要安全训练的机器人系统研发
可扩展的具身智能面临真实交互成本高、风险大的挑战。尽管具身世界模型可通过模拟推演提供前景,但现有方法存在几何幻觉且缺乏统一的策略优化框架。本文提出RoboStereo,一种对称双塔4D世界模型,通过双向跨模态增强确保时空几何一致性,缓解物理幻觉。基于此高保真4D模拟器,我们构建首个基于世界模型的统一策略优化框架:(1) 执行前验证的测试时策略增强(TTPA),(2) 利用视觉感知奖励从专家演示中学习的仿生进化策略学习(IEPL),(3) 支持自主技能发现与自我修正的开放探索策略学习(OEPL)。综合实验表明,RoboStereo生成质量达领先水平,其统一框架在细粒度操作任务上实现平均相对提升超过97%。
原文摘要 · Abstract (English)
Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models (EWMs) offer promise through imagined rollouts, existing approaches suffer from geometric hallucinations and lack unified optimization frameworks for practical policy improvement. We introduce RoboStereo, a symmetric dual-tower 4D world model that employs bidirectional cross-modal enhancement to ensure spatiotemporal geometric consistency and alleviate physics hallucinations. Building upon this high-fidelity 4D simulator, we present the first unified framework for world-model-based policy optimization: (1) Test-Time Policy Augmentation (TTPA) for pre-execution verification, (2) Imitative-Evolutionary Policy Learning (IEPL) leveraging visual perceptual rewards to learn from expert demonstrations, and (3) Open-Exploration Policy Learning (OEPL) enabling autonomous skill discovery and self-correction. Comprehensive experiments demonstrate RoboStereo achieves state-of-the-art generation quality, with our unified framework delivering >97% average relative improvement on fine-grained manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。