arXiv:2604.01666cs.CV2026-04中稿 · CVPR

用合成运动数据训练视频生成模型,提升动态场景真实感与可控性。

DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data

  • 用计算机渲染的光流数据替代真实视频训练,实现精准运动控制。
  • 在剧烈人体动作和极端镜头运动上显著提升生成质量和可控性。
  • 适合需要精细运动控制的视频生成研究者或工业应用开发者。

尽管近期取得进展,视频扩散模型在生成高度动态运动或需细粒度运动控制的视频时仍存在困难,核心瓶颈在于常用训练数据集中此类样本稀缺。为此,我们提出DynaVid,一种利用合成运动数据进行训练的视频生成框架。该方法通过计算机图形学管线生成光流表示的合成运动数据。其优势在于:一方面,合成运动提供多样化的运动模式和精确控制信号,难以从真实数据中获得;另一方面,渲染出的光流仅编码运动信息,与外观解耦,避免模型学习到不自然的合成视觉风格。基于此,DynaVid采用两阶段生成架构:先由运动生成器合成运动,再由运动引导的视频生成器根据运动生成帧。该解耦设计使模型能从合成数据中学习动态运动模式,同时保留真实视频的视觉真实性。我们在剧烈人体动作生成和极端相机运动控制两个挑战性场景中验证了该框架的有效性。大量实验表明,DynaVid在动态运动生成和相机运动控制方面显著提升了真实感与可控性。

原文摘要 · Abstract (English)

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity of such examples in commonly used training datasets. To address this, we introduce DynaVid, a video synthesis framework that leverages synthetic motion data in training, which is represented as optical flow and rendered using computer graphics pipelines. This approach offers two key advantages. First, synthetic motion offers diverse motion patterns and precise control signals that are difficult to obtain from real data. Second, unlike rendered videos with artificial appearances, rendered optical flow encodes only motion and is decoupled from appearance, thereby preventing models from reproducing the unnatural look of synthetic videos. Building on this idea, DynaVid adopts a two-stage generation framework: a motion generator first synthesizes motion, and then a motion-guided video generator produces video frames conditioned on that motion. This decoupled formulation enables the model to learn dynamic motion patterns from synthetic data while preserving visual realism from real-world videos. We validate our framework on two challenging scenarios, vigorous human motion generation and extreme camera motion control, where existing datasets are particularly limited. Extensive experiments demonstrate that DynaVid improves the realism and controllability in dynamic motion generation and camera motion control.

视频生成运动控制扩散模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。