提出可扩展的通用世界模型框架,提升仿真决策的泛化与可信度。
WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making
- 通过行为条件化缓解策略分布偏移,增强模型泛化能力。
- 引入回溯滚动机制,无需集成即可高效估计不确定性。
- 适用于少样本真实操作任务,适合机器人决策研究者使用。
世界模型在具身环境中对决策至关重要,可实现低成本模拟探索。为支持有效决策,世界模型需具备强泛化能力以应对分布外(OOD)区域,并提供可靠不确定性估计以评估模拟结果可信度,但现有可扩展方法面临挑战。本文提出WHALE框架,包含行为条件化与回溯滚动两项关键技术:行为条件化解决策略分布偏移这一主要泛化误差来源;回溯滚动实现无需模型集成的高效不确定性估计。这两项技术具有通用性,可与任意神经网络架构结合。基于此,我们构建了基于时空变换器的可扩展世界模型Whale-ST,实验表明其在价值估计准确性和视频生成保真度上表现更优。此外,我们的不确定性估计显著提升了完全离线场景下的模型基策略优化效果。我们还提出了Whale-X,一个在970K轨迹上训练的414M参数世界模型,来自Open X-Embodiment数据集。实验证明其在真实操作任务中仅需少量示范即可展现良好可扩展性与泛化能力。
原文摘要 · Abstract (English)
World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate effective decision-making, world models must be equipped with strong generalizability to support faithful imagination in out-of-distribution (OOD) regions and provide reliable uncertainty estimation to assess the credibility of the simulated experiences, both of which present significant challenges for prior scalable approaches. This paper introduces WHALE, a framework for learning generalizable world models, consisting of two key techniques: behavior-conditioning and retracing-rollout. Behavior-conditioning addresses the policy distribution shift, one of the primary sources of the world model generalization error, while retracing-rollout enables efficient uncertainty estimation without the necessity of model ensembles. These techniques are universal and can be combined with any neural network architecture for world model learning. Incorporating these two techniques, we present Whale-ST, a scalable spatial-temporal transformer-based world model with enhanced generalizability. We demonstrate the superiority of Whale-ST in simulation tasks by evaluating both value estimation accuracy and video generation fidelity. Additionally, we examine the effectiveness of our uncertainty estimation technique, which enhances model-based policy optimization in fully offline scenarios. Furthermore, we propose Whale-X, a 414M parameter world model trained on 970K trajectories from Open X-Embodiment datasets. We show that Whale-X exhibits promising scalability and strong generalizability in real-world manipulation scenarios using minimal demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。