arXiv:2607.13017cs.ROcs.CV2026-07被引 5

用光流统一表示动作,让模型更好理解视频中的运动变化。

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

论文配图:FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
图 1 · 摘自论文原文
  • 用光流作为动作的统一表示,与视频生成模型格式一致。
  • 在机器人操控任务中成功率达92.94%,世界建模准确率提升18.4%。
  • 可利用无标签视频数据预训练,适合大规模动作理解场景。

世界动作模型(WAMs)能借助预训练视频生成器实现世界建模与动作预测。然而,直接使用此类生成器进行控制带来新挑战:如何以与预训练视频生成器对齐且蕴含足够运动信息的动作表示形式进行控制。现有数值动作无法满足前者,而先前的视觉动作表示忽略帧间时序运动结构。为此,我们提出FlowWAM,一种双流扩散框架,采用光流作为统一的、视频原生的动作表示。光流视频与RGB视频格式一致,编码每像素位移信息。通过在共享预训练视频生成器中联合建模,FlowWAM可自然实现两种模式:策略模式下生成光流用于动作预测;世界建模模式下使用目标光流序列引导未来视频生成。此外,由于光流可从原始视频中无需动作标签即可提取,FlowWAM可利用大规模无标签视频数据集进行预训练。实验表明,基于光流的动作表示在两种模式下均有提升。在RoboTwin操纵任务中,清洁设置成功率提升至92.94%,随机设置达92.14%,优于VLA和WAM基线。在WorldArena世界建模任务中,整体EWMScore达到63.71,轨迹准确率相对提升18.4%。更多结果见项目网站:https://flow-wam.github.io。

原文摘要 · Abstract (English)

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control. Existing numerical actions fail to satisfy the former, and prior visual action representations overlook the temporal motion structure across frames. We address this issue with FlowWAM, a dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Flow videos share the same format as RGB videos and encode rich per-pixel displacement. By jointly modeling them within a shared pretrained video generator, FlowWAM can naturally implement two modes of WAMs. In policy mode, FlowWAM generates flow for action prediction, while in world-model mode, it uses target flow sequences to guide future video generation. Moreover, since flow can be easily extracted from raw videos without action labels, FlowWAM can leverage large-scale action-unlabeled video datasets for pretraining. We empirically find that our flow-based action representation delivers gains across both modes. On RoboTwin manipulation, FlowWAM raises the success rate to 92.94% on the Clean setting and 92.14% on Random, outperforming both VLA and WAM baselines. On WorldArena world modeling, it achieves the best overall EWMScore (63.71) with an 18.4% relative improvement in trajectory accuracy. More results can be found on our project website: https://flow-wam.github.io .

动作表示光流世界建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。