arXiv:2602.10104cs.CVcs.AI2026-02被引 6

让视频世界模型学会通用动作控制,无需标注动作。

Olaf-World: Orienting Latent Actions for Video World Modeling

  • 用动作效果差异作为锚点,对齐不同场景的潜在动作空间。
  • 在多个数据集上实现更强的零样本动作迁移与更少数据适应新控制接口。
  • 适合研究视频生成、动作可控建模与自监督学习的学者。

扩大可动作控制的世界模型受限于动作标签的稀缺性。尽管潜在动作学习能从无标签视频中提取控制接口,但学习到的潜在变量往往无法跨上下文迁移:它们混杂了场景特异性线索,且缺乏共享坐标系。这是因为标准目标仅在每个视频片段内运作,无法跨上下文对齐动作语义。我们的关键洞察是:尽管动作不可观测,但其语义效应可被观测,并可用作共享参考。我们提出Seq$Δ$-REPA,一种序列级控制效果对齐目标,将集成潜在动作锚定于冻结的自监督视频编码器所产生的时序特征差异。基于此,我们构建Olaf-World,一个从大规模被动视频预训练动作条件视频世界模型的流程。大量实验表明,该方法学习到更结构化的潜在动作空间,相比现有最优基线,在零样本动作迁移和新控制接口的数据高效适应方面表现更优。

原文摘要 · Abstract (English)

Scaling action-controllable world models is limited by the scarcity of action labels. While latent action learning promises to extract control interfaces from unlabeled video, learned latents often fail to transfer across contexts: they entangle scene-specific cues and lack a shared coordinate system. This occurs because standard objectives operate only within each clip, providing no mechanism to align action semantics across contexts. Our key insight is that although actions are unobserved, their semantic effects are observable and can serve as a shared reference. We introduce Seq$Δ$-REPA, a sequence-level control-effect alignment objective that anchors integrated latent action to temporal feature differences from a frozen, self-supervised video encoder. Building on this, we present Olaf-World, a pipeline that pretrains action-conditioned video world models from large-scale passive video. Extensive experiments demonstrate that our method learns a more structured latent action space, leading to stronger zero-shot action transfer and more data-efficient adaptation to new control interfaces than state-of-the-art baselines.

视频建模潜在动作自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。