用深度信息提升机器人真实户外数据的世界模型泛化能力
Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data

- 引入深度作为几何先验,结合正则化约束潜在表示多样性
- 在农业机器人视频上训练,视觉里程计误差降33%,长程预测更稳定
- 轻量级设计不增加推理开销,适合实际部署的机器人系统
世界模型,尤其是基于JEPA架构的模型,已被证明能学习多种环境的鲁棒动态。然而,从视觉复杂的现实世界数据中学习仍具挑战性,尤其在不可预测的户外环境中。本文在训练中引入深度作为几何先验,直接从机器人视频数据中学习更鲁棒的潜在动态,并应对视觉复杂性。该方法结合深度监督与各向同性潜在正则化(SIGReg),在最大化任务无关潜在多样性的同时,约束其组织方式,目标是获得与场景几何一致的最高熵表示。为在不增加推理时间的前提下满足更高复杂度,还引入仅训练时使用的过参数化。在真实农业机器人视频上训练一个1800万参数模型,通过冻结表示的视觉里程计探针、基于预测的意外检测以及多步潜在滚动保真度进行评估。相比基线LeWM,本方法将视觉里程计探针误差降低33%,显著提升域内和域外TartanGround基准上的意外分数分离度,并在领域迁移下改善多步滚动保真度,且收益随滚动时长增加。值得注意的是,对光照、阴影等非三维几何相关的物理理解也表现出意外分数分离的提升。结果表明,轻量级训练时几何先验使紧凑的JEPA世界模型在真实户外数据上更具实用性与可迁移性,同时保持无推理开销。工作表明,深度作为物理基础先验可增强世界模型在各类任务上的泛化能力。
原文摘要 · Abstract (English)
World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments. However, learning from visually complex real-world data remains a challenge, especially in unpredictable outdoor environments. We introduce depth as a geometric prior during training in learning more robust latent dynamics directly from robot video data and handling visual complexity. This combines depth supervision with an isotropy-inducing latent regularizer (SIGReg), maximizing task-agnostic latent diversity while constraining how that diversity is organized, with the combined objective targeting the highest-entropy representation consistent with scene geometry. To satisfy this greater complexity without increasing inference time, we also add training-only overparameterization. Training an 18M-parameter model on video from a real agricultural robot, we evaluate with frozen-representation visual odometry probes, predictor-based surprise detection, and multi-step latent rollout fidelity. Compared to the baseline LeWM, our method lowers visual odometry probe error by 33%, substantially increases surprise-score separation both in-domain and on the out-of-domain TartanGround benchmark, and improves multi-step rollout fidelity under domain shift, with gains that grow with rollout horizon. Notably, we also see improvements in surprise-score separation on physics understanding that is not directly tied to 3D geometry, such as lighting and shadows. These results show that a lightweight training-time geometric prior makes a compact JEPA world model more useful and more transferable on real outdoor data with strong underlying representations, without adding inference overhead. Our work suggests that depth as a physically grounded prior can enhance world model generalization on a variety of tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。