改进世界模型的表征,让机器人更会抓取和操作
Temporally Centered SIGReg Improves LeWorldModel Representations for Robot Policy Learning

- 对时序中心残差施加正则化,分离长期与短期变化
- 在LIBERO上政策成功率提升1.66倍,平均成功率从63.6%升至83.8%
- 无需预训练即可超越扩散策略和OpenVLA基线,适合机器人控制
近期的LeWorldModel(LeWM)研究表明,草图各向同性高斯正则化(SIGReg)通过将潜在表示正则化为各向同性高斯分布,实现了从像素端到端的世界模型稳定学习。然而,原始的LeWM表示对下游机器人策略学习效果不佳。本文通过蒙特卡洛分析发现,原始LeWM目标会将方差分配偏向时间持续成分,从而抑制时序中心残差的方差。实验验证,这导致残差变化被压制,机器人状态与动力学(尤其是夹爪动力学)解码能力下降。为此,我们提出将SIGReg应用于时序中心残差而非整个潜在表示。该方法解耦了持久成分与残差成分的方差分配,同时保持防坍缩特性。在LIBERO基准上,该方法使目标套件的下游策略成功率提升1.66倍,所有套件平均成功率从63.6%提升至83.8%。不依赖外部预训练的情况下,其性能优于从零训练的扩散策略和预训练的OpenVLA基线。结果表明,原始LeWM的方差分配偏差是导致下游策略差距的原因,解耦持久与残差变化可获得更适合机器人策略学习的表示。
原文摘要 · Abstract (English)
Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world model learning from pixels by regularizing the latent representation toward an isotropic Gaussian. While effective for latent-space planning, the representations learned by Raw LeWM are poorly suited for downstream robot policy learning. In this paper, through Monte Carlo analysis, we show that the Raw LeWM objective biases variance allocation toward the temporally persistent component, thereby suppressing the variance of the temporally centered residual. Consistent with this analysis, trained Raw LeWM representations exhibit suppressed residual variation and reduced decodability of robot state and dynamics, particularly gripper dynamics, which are crucial for robotic manipulation. To address this issue, we apply SIGReg to temporally centered residuals rather than to the whole latent representation. This simple change decouples persistent and residual variance allocation while retaining an effective anti-collapse property. On the LIBERO benchmark, our method improves downstream policy success on the Goal suite by 1.66x and raises the average success rate across all suites from 63.6% to 83.8%. Without external pretraining, it also outperforms both Diffusion Policy trained from scratch and the pretrained OpenVLA baseline. These results associate the variance-allocation bias of Raw LeWM with the downstream policy gap, and show that decoupling persistent and residual variation yields representations better suited for downstream robot policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。