arXiv:2509.15130cs.GRcs.AI2025-09中稿 · CVPR被引 14

无需训练即可精准控制视频生成的3D/4D相机运动,保持画面真实感。

Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control

  • 推理时通过三阶段机制解耦镜头与场景,实现零样本相机控制
  • 在多个基准上达到最优轨迹一致性与视觉保真度,超越有训练和无训练基线
  • 适用于视频编辑、虚拟试穿等十余种下游任务,通用性强

视频扩散模型具备丰富的世界先验,但在空间任务中受限于控制能力差、时空不一致及场景-相机动态耦合。现有方法如特定任务微调或后处理扭曲常引入视觉伪影、泛化性差或计算开销高。我们提出WorldForge,一种纯推理时、无需训练的新框架,解决上述问题。其包含三个协同组件:首先,步骤内迭代修正环在去噪过程中注入细粒度运动引导,确保输出严格遵循目标相机路径;其次,基于光流分析识别并分离潜在空间中的运动通道,选择性施加引导,解耦运动与外观,保留视觉质量;第三,双路径引导策略通过对比有引导与无引导的去噪路径,自适应纠正漂移,有效消除结构输入错位引发的伪影。三者结合,在不重训练模型的前提下实现精准轨迹对齐与照片级合成。作为即插即用、模型无关的解决方案,WorldForge展现出高度泛化能力,不仅支持鲁棒的零样本3D/4D生成,还可无缝赋能十余种下游应用,如视频编辑、稳定化与虚拟试穿。大量实验验证其在轨迹遵循性和感知质量上均达当前最优,优于依赖训练与仅推理的基线方法。

原文摘要 · Abstract (English)

Video diffusion models have rich world priors, but their use in spatial tasks is limited by poor control, spatial-temporal inconsistent results, and entangled scene-camera dynamics. Current approaches, such as per-task fine-tuning or post-process warping, often introduce visual artifacts, fail to generalize, or incur high computational costs. We introduce WorldForge, a novel, training-free framework that operates purely at inference time to resolve these issues. Our method comprises three synergistic components. First, an intra-step refinement loop injects fine-grained motion guidance during the denoising process, iteratively correcting the output to ensure strict adherence to the target camera path. Second, an optical flow-based analysis identifies and isolates motion-related channels within the latent space. This allows our framework to selectively apply guidance, thereby decoupling motion from appearance and preserving visual fidelity. Third, a dual-path guidance strategy adaptively corrects for drift by comparing the guided generation against an unguided, reference denoising path, effectively neutralizing artifacts caused by misaligned structural inputs. Together, these components inject precise, trajectory-aligned control without model retraining, achieving accurate motion guidance and photorealistic synthesis. As a plug-and-play, model-agnostic solution, WorldForge demonstrates highly versatile generalizability. Beyond robust zero-shot 3D/4D generation, it readily empowers over a dozen diverse downstream applications, seamlessly enabling tasks like video editing, stabilization, and virtual try-on. Extensive experiments confirm state-of-the-art performance in trajectory adherence and perceptual quality, outperforming both training-dependent and inference-only baselines.

视频生成扩散模型相机控制零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。