通过重塑优化景观提升稀疏视角下相机位姿估计精度。
Landscape-Awareness for Geometric View Diffusion Model

- 用基于得分的方法重构优化空间,引导梯度更新
- 在双视角条件下收敛更快,样本效率更高
- 适合需要高精度与低计算开销的3D重建任务
稀疏视角下的准确相机位姿估计仍是挑战,尤其在双视角场景中。现有方法借助扩散模型(如Zero123)在相对位姿条件下合成新视图,通过MSE损失优化实现位姿估计,但常受非凸损失景观和大量局部极小值影响,对初始化敏感且依赖简单的多起点策略。我们分析了这些优化难题并可视化失败案例,发现几何模糊性(如对称性和自相似性)会误导梯度更新至错误位姿。为此,提出一种基于得分的方法重塑优化景观,引导更新趋向真实位姿,随后使用位姿条件扩散模型进行精修。实验表明,该方法提升收敛性,减少对暴力采样的依赖,并在保持竞争性精度的同时显著提高样本效率。
原文摘要 · Abstract (English)
Accurate camera viewpoint estimation under sparse-view conditions remains challenging, particularly in two-view scenarios. Recent approaches leverage diffusion models such as Zero123 to synthesize novel views conditioned on relative viewpoint, showing promising results when repurposed for viewpoint estimation via optimization with MSE loss. However, existing methods often suffer from nonconvex loss landscape with numerous local minima, making them sensitive to initialization and reliant on naive multistart strategies. We analyze these optimization challenges and visualize failure cases, showing that geometric ambiguities, such as symmetry and self-similarity, can mislead gradient-based updates toward incorrect viewpoints. To address these limitations, we propose a score-based method that reshapes the optimization landscape to guide updates toward the ground-truth viewpoint, followed by a refinement stage using a viewpoint-conditioned diffusion model. Experiments show that our method improves convergence, reduces reliance on brute-force sampling, and achieves competitive accuracy with higher sample-efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。