arXiv:2604.09330cs.ROcs.CV2026-04被引 4

用统一框架同步生成视频与动作,解决机器人数据合成难题。

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

  • 双流结构联合生成视频与动作,视觉语言条件驱动。
  • 生成视频与动作对齐度高,支持可执行轨迹回放。
  • 适合需要高质量合成数据的机器人强化学习研究者。

近期基于大规模人类远程操控数据训练的机器人基础模型已使机器人能完成复杂现实任务。但系统扩展困难,因特定任务示范收集成本高且耗时。合成数据(尤其是生成视频)提供了可行方向,但现有世界模型(WMs)不提供配对动作轨迹,难以用于策略学习。世界-动作(WA)模型虽部分解决此问题,但常存在视频与动作对齐不足;两阶段生成先出视频再推断动作,则效率低且误差累积。为此,我们提出VAG——一种基于流匹配的双流统一框架,在视觉与语言条件下联合生成视频与动作。通过同步双分支去噪,并使用自适应3D池化机制将紧凑全局视频上下文传递至动作分支,提升生成过程中的跨模态一致性。在模拟与真实场景中,VAG均生成高质量对齐的视频-动作对,具备良好预测性能,支持可执行轨迹回放,并可作为有用合成预训练数据,提升下游策略泛化能力,展现出作为具身数据合成实用型世界-动作模型的潜力。

原文摘要 · Abstract (English)

Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-world tasks. However, scaling these systems remains difficult because collecting task-specific demonstrations is expensive and labor-intensive. Synthetic data, especially generated videos, offer a promising direction, but existing World Models (WMs) are not directly suitable for policy learning since they do not provide paired action trajectories. World-Action (WA) models partially address this by predicting actions with visual outputs, yet often lack strong video-action alignment, while two-stage pipelines that generate video first and then infer actions introduce inefficiency and error accumulation. To address these limitations, we propose VAG, a unified flow-matching-based dual-stream framework that jointly generates video and action under visual and language conditioning. By synchronizing denoising in both branches and using an adaptive 3D pooling mechanism to transfer compact global video context to the action branch, VAG improves cross-modal consistency during generation. Across both simulated and real-world settings, VAG produces aligned video-action pairs with competitive prediction quality, supports executable trajectory replay, and provides useful synthetic pretraining data that improves downstream policy generalization, indicating its potential as a practical world-action model for embodied data synthesis.

视频生成机器人合成数据动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。