arXiv:2510.27545cs.ROcs.AI2025-10被引 2

用能量模型提升机器人策略的稳定性与推理能力,实现快速收敛与零样本纠错。

EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities

  • 基于能量模型构建新策略架构,通过能量函数实现稳定推理
  • 部分任务仅需2步推理即收敛,相较扩散模型提速50倍
  • 无需额外训练即可零样本恢复失败动作,适合真实场景部署

以生成模型参数化的隐式策略(如Diffusion Policy)已成为机器人视觉-语言-动作模型的标准,但普遍存在计算成本高、暴露偏差和推断不稳定问题,导致分布偏移下易发散。能量模型(EBMs)通过端到端学习能量势场并建模平衡动力学,可提升鲁棒性并减少暴露偏差,但其策略应用长期难以有效扩展。近期能量基变换器(EBTs)证明了EBMs在高维空间的可扩展性,但在物理具身模型中的潜力仍待挖掘。本文提出新型能量基策略架构EBT-Policy,解决了机器人与真实世界任务中的核心挑战。在仿真与真实任务中,EBT-Policy持续优于扩散基策略,且训练与推理开销更低。令人惊讶的是,在某些任务中仅需2步推理即可收敛,相比Diffusion Policy的100步降低50倍。此外,该模型展现出此前未见的涌现能力:仅通过行为克隆,即可在不进行显式重试训练的情况下,实现对失败动作序列的零样本恢复。通过利用标量能量实现不确定性感知推理与动态计算分配,EBT-Policy为应对分布偏移下的鲁棒、泛化机器人行为提供了可行路径。

原文摘要 · Abstract (English)

Implicit policies parameterized by generative models, such as Diffusion Policy, have become the standard for policy learning and Vision-Language-Action (VLA) models in robotics. However, these approaches often suffer from high computational cost, exposure bias, and unstable inference dynamics, which lead to divergence under distribution shifts. Energy-Based Models (EBMs) address these issues by learning energy landscapes end-to-end and modeling equilibrium dynamics, offering improved robustness and reduced exposure bias. Yet, policies parameterized by EBMs have historically struggled to scale effectively. Recent work on Energy-Based Transformers (EBTs) demonstrates the scalability of EBMs to high-dimensional spaces, but their potential for solving core challenges in physically embodied models remains underexplored. We introduce a new energy-based architecture, EBT-Policy, that solves core issues in robotic and real-world settings. Across simulated and real-world tasks, EBT-Policy consistently outperforms diffusion-based policies, while requiring less training and inference computation. Remarkably, on some tasks it converges within just two inference steps, a 50x reduction compared to Diffusion Policy's 100. Moreover, EBT-Policy exhibits emergent capabilities not seen in prior models, such as zero-shot recovery from failed action sequences using only behavior cloning and without explicit retry training. By leveraging its scalar energy for uncertainty-aware inference and dynamic compute allocation, EBT-Policy offers a promising path toward robust, generalizable robot behavior under distribution shifts.

机器人策略能量模型零样本恢复高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。