arXiv:2606.20104cs.LGcs.AI2026-06被引 4

让模型通过动作反推感知,学出更可控的环境表示。

Sensorimotor World Models: Perception for Action via Inverse Dynamics

论文配图:Sensorimotor World Models: Perception for Action via Inverse Dynamics
图 1 · 摘自论文原文
  • 用逆动力学约束训练潜空间,避免表征坍塌。
  • 在无奖励离线数据上实现稳定建模与高效规划。
  • 适合做端到端动作感知建模的研究者参考。

感知为行动服务,意味着世界表征应不仅追求视觉保真度,更要对行动相关。现有的隐式JEPA型世界模型虽能从高维观测中学习紧凑的预测状态,但端到端训练困难,因仅追求可预测性易导致表征坍塌。本文提出传感器运动世界模型(SMWM):一种通过逆动力学正则化实现端到端训练的潜在世界模型。该正则化同时解决两个问题:防止表征坍塌,并生成与动作对齐的表征。通过强制潜态保留过渡背后动作的信息,模型聚焦于环境的可控制自由度,过滤不可控干扰。由此,在无需冻结编码器、指数移动平均或复杂潜空间正则化的情况下,仅用离线、无奖励轨迹即可稳定训练。实验表明,SMWM学习到紧凑且可解释的潜空间,在2D和3D简单控制任务中达到竞争性规划性能。

原文摘要 · Abstract (English)

Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions. At the same time, latent JEPA-style world models advocate learning compact predictive states from high-dimensional observations to facilitate the prediction of future states, but end-to-end training of these models is nontrivial because representations may collapse if our only goal is to construct a latent state that is easy to predict. We introduce a sensorimotor world model (SMWM): a latent world model trained end-to-end with inverse dynamics regularization. This single regularizer addresses both issues: it prevents representation collapse and induces action-aligned representations. By forcing latent states to preserve information about the action underlying a transition, it biases the model toward the controllable degrees of freedom of the environment while discarding uncontrollable distractors. This yields stable latent world models trained from offline, reward-free trajectories, without frozen encoders, exponential moving averages, or complex latent regularizers. Empirically, SMWM learns compact, interpretable latent spaces and enables competitive planning performance across simple 2D and 3D control tasks.

世界模型动作感知逆动力学潜空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。