arXiv:2503.06508cs.CVcs.AI2025-03被引 7

无需微调即可模拟相机运动,让视频生成更轻量高效

LightMotion: A Light and Tuning-free Method for Simulating Camera Motion in Video Generation

  • 在潜在空间通过置换与重采样模拟平移、缩放、旋转等运动
  • 引入噪声修正机制,解决潜空间信噪比变化导致的画质下降
  • 适合追求快速生成且不想复杂调参的视频生成研究者

现有控制相机运动的视频生成方法存在微调和推理计算瓶颈。本文提出LightMotion,一种轻量且无需微调的视频生成相机运动模拟方法。该方法在潜在空间操作,避免了额外微调、图像修复和深度估计,流程更简洁。具体贡献包括:(i) 潜在空间置换操作可有效模拟平移、缩放、旋转等多种相机运动;(ii) 潜在空间重采样策略结合背景感知采样与跨帧对齐,准确填补新视角并保持帧间一致性;(iii) 深入分析发现,置换与重采样会引发潜在空间信噪比(SNR)偏移,导致生成质量下降。为此,提出潜在空间校正,通过在去噪过程中重新引入噪声以缓解SNR偏移,提升生成质量。大量实验表明,LightMotion在定量和定性指标上均优于现有方法。

原文摘要 · Abstract (English)

Existing camera motion-controlled video generation methods face computational bottlenecks in fine-tuning and inference. This paper proposes LightMotion, a light and tuning-free method for simulating camera motion in video generation. Operating in the latent space, it eliminates additional fine-tuning, inpainting, and depth estimation, making it more streamlined than existing methods. The endeavors of this paper comprise: (i) The latent space permutation operation effectively simulates various camera motions like panning, zooming, and rotation. (ii) The latent space resampling strategy combines background-aware sampling and cross-frame alignment to accurately fill new perspectives while maintaining coherence across frames. (iii) Our in-depth analysis shows that the permutation and resampling cause an SNR shift in latent space, leading to poor-quality generation. To address this, we propose latent space correction, which reintroduces noise during denoising to mitigate SNR shift and enhance video generation quality. Exhaustive experiments show that our LightMotion outperforms existing methods, both quantitatively and qualitatively.

视频生成相机运动扩散模型轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。