用动作场建模机器人视觉,让生成视频更精准反映机械臂运动轨迹。
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields

- 将动作直接映射为视觉空间中的结构化动作场,实现几何对齐。
- 在WorldArena上超越现有基线,显著提升生成视频的几何保真度。
- 适合需要精确物理交互模拟的机器人视觉建模任务。
预训练视频扩散模型提供了强大的时空生成先验,是构建机器人世界模型的理想基础。尽管近期世界-动作模型联合优化未来视频与动作,但多数将视频生成视为策略学习的辅助表示,未能充分探索逆问题:利用动作信号引导视频合成,导致生成轨迹中常丢失机器人空间几何精度和精细的机器人-物体交互动态。为此,我们提出EA-WM,一种事件感知的生成式世界模型,有效打通运动控制与视觉感知的闭环。不同于将关节或末端执行器动作作为抽象低维标记注入,EA-WM将动作与运动状态直接投影至目标相机视图,形成结构化的运动-视觉动作场。为充分利用这一几何根基表示,我们引入事件感知双向融合模块,调节跨分支注意力,捕捉物体状态变化与交互动态。在综合性WorldArena基准上的评估显示,EA-WM达到当前最优性能,显著优于现有基线。
原文摘要 · Abstract (English)
Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly treat video generation as an auxiliary representation for policy learning. Consequently, they insufficiently explore the inverse problem: leveraging action signals to guide video synthesis, thereby often failing to preserve precise robot spatial geometry and fine-grained robot-object interaction dynamics in the generated rollouts. To bridge this gap, we present EA-WM, an Event-Aware Generative World Model that effectively closes the loop between kinematic control and visual perception. Rather than injecting joint or end-effector actions as abstract, low-dimensional tokens, EA-WM projects actions and kinematic states directly into the target camera view as Structured Kinematic-to-Visual Action Fields. To fully exploit this geometrically grounded representation, we introduce event-aware bidirectional fusion blocks that modulate cross-branch attention, capturing object state changes and interaction dynamics. Evaluated on the comprehensive WorldArena benchmark, EA-WM achieves state-of-the-art performance, outperforming existing baselines by a significant margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。