arXiv:2601.05230cs.AIcs.CV2026-01被引 41

从真实视频中学习动作潜空间,实现无需标签的环境预测。

Learning Latent Action World Models In The Wild

  • 仅用视频训练连续受限的潜动作模型,避免依赖动作标签。
  • 能捕捉人类进入房间等复杂环境变化,跨视频泛化动作语义。
  • 潜动作可作为通用接口,支持规划任务且性能媲美有标签模型。

具备现实世界推理与规划能力的智能体需预测行为后果。虽然世界模型具备此能力,但通常依赖动作标签,而大规模获取标签成本高。这促使了潜动作模型的发展,即仅从视频中学习动作空间。本文研究在真实场景视频上学习潜动作世界模型,扩展了以往局限于简单机器人仿真、游戏或操作数据的研究范围。尽管真实视频包含更多样化的动作,也带来了环境噪声、缺乏统一身体形态等挑战。我们探讨了动作应具备的特性、架构选择及评估方法。发现连续但受约束的潜动作能有效捕捉真实视频中的复杂动作,优于常见的向量量化方法。例如,人类进入房间等环境变化可跨视频传递。在无统一身体形态的情况下,潜动作主要在相机视角下局部化。但仍可训练控制器将已知动作映射到潜动作,使潜动作成为通用接口,并在规划任务中达到与条件动作基线相当的性能。分析与实验为潜动作模型向真实世界扩展迈出一步。

原文摘要 · Abstract (English)

Agents capable of reasoning and planning in the real world require the ability of predicting the consequences of their actions. While world models possess this capability, they most often require action labels, that can be complex to obtain at scale. This motivates the learning of latent action models, that can learn an action space from videos alone. Our work addresses the problem of learning latent actions world models on in-the-wild videos, expanding the scope of existing works that focus on simple robotics simulations, video games, or manipulation data. While this allows us to capture richer actions, it also introduces challenges stemming from the video diversity, such as environmental noise, or the lack of a common embodiment across videos. To address some of the challenges, we discuss properties that actions should follow as well as relevant architectural choices and evaluations. We find that continuous, but constrained, latent actions are able to capture the complexity of actions from in-the-wild videos, something that the common vector quantization does not. We for example find that changes in the environment coming from agents, such as humans entering the room, can be transferred across videos. This highlights the capability of learning actions that are specific to in-the-wild videos. In the absence of a common embodiment across videos, we are mainly able to learn latent actions that become localized in space, relative to the camera. Nonetheless, we are able to train a controller that maps known actions to latent ones, allowing us to use latent actions as a universal interface and solve planning tasks with our world model with similar performance as action-conditioned baselines. Our analyses and experiments provide a step towards scaling latent action models to the real world.

潜空间世界模型真实视频动作学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。