arXiv:2607.03732cs.CV2026-07

用粗略视频控制生成精细动作,无需训练即可实现动态可控的视频生成。

ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics

论文配图:ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
图 1 · 摘自论文原文
  • 以物理模拟或真实录制的粗略视频作为动力学载体,引导生成过程。
  • 在不依赖配对数据的情况下,保持动作真实性并匹配文本描述。
  • 适合需要精准运动控制的视频生成场景,如动画制作与虚拟仿真。

现代视频生成模型难以精确控制复杂动态,仅靠文本提示常无法指定物理上合理的细粒度运动与交互。本文提出“代理条件视频生成”,利用基于物理的模拟或真实记录的粗略代理视频作为动力学载体,控制前景物体的运动。给定代理视频和文本提示,目标是合成一个保留代理动态、生成新内容且符合提示的合理交互的新视频。由于成对的代理-目标视频难以获取,我们提出无训练框架ProxyUp,基于预训练视频生成模型构建。ProxyUp首先将代理视频反演为中间潜在表示,并应用“区域级潜在噪声注入”,保留运动关键的代理潜在变量,同时在需文本驱动重生成的区域注入噪声。为缓解此启发式潜在组合带来的分布偏差与弱前景-背景耦合问题,进一步提出“随机流松弛(SFR)”,在欧拉积分采样前逐步将组合潜在表示向模型学习的分布逼近。在模拟和真实代理上的实验表明,ProxyUp在动态保真度与文本对齐方面优于强基线的视频编辑与运动迁移方法。

原文摘要 · Abstract (English)

Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plausible, fine-grained motion and interactions. We introduce $\textit{proxy-conditioned video generation}$, where a coarse proxy video from physics-based simulation or real-world recording serves as a dynamics carrier to control foreground object motion. Given a proxy video and a text prompt, the goal is to synthesize a new video that preserves the proxy dynamics while generating novel content and plausible interactions aligned with the prompt. Since paired proxy-target videos are difficult to obtain, we propose $\textbf{ProxyUp}$, a training-free framework built on pretrained video generative models. ProxyUp first inverts the proxy video into an intermediate latent representation and applies $\textbf{region-wise latent noising}$, preserving motion-critical proxy latents while injecting noise into regions intended for text-driven regeneration. To mitigate the distribution mismatch and weak foreground-background coupling introduced by this heuristic latent composition, we further propose $\textbf{Stochastic Flow Relaxation (SFR)}$, which progressively relaxes the composed latent toward the model's learned distribution before ODE sampling. Experiments on both simulation and real-world proxies show that ProxyUp outperforms strong video editing and motion transfer baselines in dynamic fidelity and text alignment.

视频生成动态控制无训练代理视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。