arXiv:2609.03639cs.CV2026-09

通过分步控制相机运动,让单图生成视频更稳定。

Stabilizing Camera-Controlled Novel View Synthesis at Inference Time

论文配图:Stabilizing Camera-Controlled Novel View Synthesis at Inference Time
图 1 · 摘自论文原文
  • 将相机运动拆成小步自回归生成,减少每步畸变
  • 每步运动≤18-20°时效果稳定,超此范围性能下降
  • 适合需要长序列生成的3D重建与视觉应用

使用预训练视频扩散模型进行无需训练的相机控制新视角合成,在大相机运动和长生成序列下常不稳定。现有方法通常组合多个推理阶段组件,难以判断关键设计。我们发现稳定性主要源于简单机制:将相机运动分解为小步自回归生成,可限制每步几何畸变并减少误差累积。控制相机步长实验表明,当每步运动接近18-20°时性能开始显著下降。我们进一步评估了几何约束空间注意力、低频外观锚定作为辅助优化,并采用高效无注册的变形流水线。在RealEstate10K和MegaScene数据集上,CamTrol++相比无需训练基线,提升了时间与几何一致性、下游3D重建质量及生成效率,支持56帧生成且对可控深度噪声具有鲁棒性。结果表明,仅在推理时精细控制相机运动即可显著提升稳定性,无需重训练或修改扩散主干。

原文摘要 · Abstract (English)

Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion models often becomes unstable under large camera motion and long generation horizons. Existing approaches commonly combine several inference-time components, making it unclear which design choices are most important for stability. We show that the main source of stability is simple. Decomposing camera motion into small autoregressive steps limits per-step geometric distortion and reduces error accumulation. A controlled camera-step study shows that performance remains stable for small motions and degrades more strongly as the per-step motion approaches $18$-$20^\circ$. We further evaluate geometry-constrained spatial attention and low-frequency appearance anchoring as supporting refinements, together with an efficient registration-free warping pipeline. Across RealEstate10K and MegaScene, CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines. The method remains effective for 56-frame generation and under substantial controlled depth corruption. These results show that careful control of camera motion at inference time can substantially improve the stability of camera-controlled novel view synthesis without retraining or modifying the diffusion backbone.

新视角合成扩散模型相机控制生成稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。