arXiv:2606.04130cs.RO2026-06

无需动作标签,从视频中学习可解释的连续动作表示。

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization

论文配图:CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization
图 1 · 摘自论文原文
  • 通过对抗性隐变量正则化与扩散生成,端到端学习动作-环境动态
  • 在无标注视频上实现行为克隆与目标导向规划,效果优于现有方法
  • 适合需要零样本动作推理与视觉导航的研究者

我们提出CLAW,一种完全端到端的自监督框架,可直接从无动作视频中联合学习世界模型与连续隐变量动作表示。该方法利用对抗性隐变量正则化与基于扩散的视频生成,捕捉结构化且语义清晰的动作表征,同时建模丰富的预测性环境动态,无需任何动作标签或注释。通过同步训练隐动作模型与世界模型,CLAW仅凭视觉观测即可推断动作如何引发环境状态转移。实验表明,所学隐动作世界模型支持仅从观察进行模仿学习和目标导向规划:从原始视频中提取的隐动作可用于行为克隆;而生成的隐动作序列可映射为可执行动作以达成目标。在多种任务与机器人形态上的大量实验验证了其有效性,结果表明该模型能生成语义合理的隐动作表示,支持高效动作迁移,并实现从观察出发的规划与模仿,性能超越现有方法。

原文摘要 · Abstract (English)

We introduce CLAW, a fully end-to-end self-supervised framework for learning a world model jointly with continuous latent action representations directly from action-free videos. Our approach leverages adversarial latent regularization and diffusion-based video generation to capture structured and semantically meaningful action representations while modeling rich, predictive environment dynamics, without relying on any action labels or annotations. By simultaneously training the Latent Action Model and world model, CLAW learns to reason about how inferred actions induce environment transitions from visual observations alone. We show that the resulting latent action world model supports both imitation learning from observation and goal-directed planning. In imitation learning, latent actions extracted from raw videos enable behavior cloning. For planning, CLAW generates sequences of latent actions and maps them to executable actions to reach desired goals. Extensive experiments across diverse tasks and embodiments demonstrate that CLAW produces semantically meaningful latent action representations, supports effective action transfer, and enables planning and imitation from observation, outperforming existing methods.

世界模型动作表示自监督扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。