arXiv:2605.30774cs.CV2026-05被引 2

让视频生成中的相机运动更真实,通过噪声空间直接控制镜头轨迹。

CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

论文配图:CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping
图 1 · 摘自论文原文
  • 将相机位姿直接嵌入噪声空间,避免画面失真
  • 利用几何重投影流保持运动轨迹一致性
  • 适合需要精准镜头控制的视频生成场景

精确的相机姿态控制对视频扩散模型至关重要,但保持几何一致性仍具挑战。现有方法直接将数值相机参数注入扩散主干网络,难以弥合抽象坐标与视觉内容之间的差距,导致结构失真。为此,我们提出CameraNoise,一种基于几何-光流引导的噪声扭曲方法,将相机运动编码为时间连贯的随机表示。不同于传统条件输入,CameraNoise直接将相机姿态嵌入噪声空间,实现运动与场景外观的解耦,同时忠实保留轨迹动态。具体而言,我们引入新型几何引导重投影流与噪声扭曲算法,联合保证扩散过程的高斯先验,并确保相机变换下的噪声传播一致性。将CameraNoise集成至扩散流程后,框架可生成稳定且高保真的视频。大量实验表明,该方法在视觉质量与轨迹忠实度上均显著优于先前方法。项目页面与代码见:https://gulucaptain.github.io/CameraNoise/。

原文摘要 · Abstract (English)

Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly inject numerical camera parameters into the diffusion backbone often fail to bridge the gap between abstract coordinates and visual content, leading to structural distortions. To address this issue, we propose CameraNoise, a flow-to-noise warping method that encodes camera motion into a temporally coherent stochastic representation. Unlike conventional conditioning, CameraNoise embeds camera poses directly into the noise space. This decouples motion from scene appearance while faithfully preserving trajectory dynamics. Specifically, we introduce a novel Geometry-guided Reprojection Flow and a noise warping algorithm, which jointly preserve the Gaussian prior of diffusion and ensure consistent noise propagation under camera transformations. By integrating CameraNoise into the diffusion process, our framework delivers stable, high-fidelity videos. Extensive experiments demonstrate that our approach significantly outperforms prior methods in both visual quality and trajectory faithfulness. The project page and code are available at: https://gulucaptain.github.io/CameraNoise/.

视频生成扩散模型相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。