通过置信度加权实现视频扩散中精准相机控制,重建大视角变化下的未见区域。
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
- 用置信度加权的点云潜变量与噪声结合,初始化扩散过程。
- 引入卡尔曼式预测-更新机制,稳定跟随预设相机轨迹。
- 适合需要高几何一致性视频生成的研究者或工业应用。
针对仅凭两张输入图像在大幅视角变化下进行新视图合成的挑战,现有基于回归的方法无法重建未见区域,而相机引导的扩散模型常因点云投影噪声或相机位姿条件不足导致轨迹偏离。为此,我们提出ConfCtrl,一种置信度感知的视频插值框架,使扩散模型在遵循指定相机位姿的同时完成未见区域重建。ConfCtrl通过将置信度加权的投影点云潜变量与噪声结合作为条件输入来初始化扩散过程,并采用类卡尔曼滤波的预测-更新机制:将投影点云视为含噪观测,利用学习到的残差修正平衡位姿驱动预测与噪声几何观测。该方法能依赖可靠投影,降低不确定区域的影响,实现稳定且几何感知的生成。多数据集实验表明,ConfCtrl生成的视图具有几何一致性且视觉合理,有效重建了大视角变化下的遮挡区域。
原文摘要 · Abstract (English)
We address the challenge of novel view synthesis from only two input images under large viewpoint changes. Existing regression-based methods lack the capacity to reconstruct unseen regions, while camera-guided diffusion models often deviate from intended trajectories due to noisy point cloud projections or insufficient conditioning from camera poses. To address these issues, we propose ConfCtrl, a confidence-aware video interpolation framework that enables diffusion models to follow prescribed camera poses while completing unseen regions. ConfCtrl initializes the diffusion process by combining a confidence-weighted projected point cloud latent with noise as the conditioning input. It then applies a Kalman-inspired predict-update mechanism, treating the projected point cloud as a noisy measurement and using learned residual corrections to balance pose-driven predictions with noisy geometric observations. This allows the model to rely on reliable projections while down-weighting uncertain regions, yielding stable, geometry-aware generation. Experiments on multiple datasets show that ConfCtrl produces geometrically consistent and visually plausible novel views, effectively reconstructing occluded regions under large viewpoint changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。