arXiv:2607.15912cs.CV2026-07

融合相对位移与特征轨迹,提升全局三维重建精度与鲁棒性

HETA++: Global Structure-from-Motion with Hybrid Explicit Translation Averaging

论文配图:HETA++: Global Structure-from-Motion with Hybrid Explicit Translation Averaging
图 1 · 摘自论文原文
  • 结合相对位移和特征轨迹进行联合优化
  • 在多个真实数据集上优于当前最优方法,误差更小、效率更高
  • 适合需要高精度三维重建的视觉定位任务

全局结构光恢复(SfM)相比增量式方法在效率和误差分布上更具优势,但平移平均仍具挑战。现有方法多依赖相对位移或特征轨迹,在相机共线运动下性能下降,且易受异常值影响。本文提出一种新型混合显式平移平均框架,同时利用相对位移和特征轨迹。首先,基于全局相机旋转精炼相对位移,并剔除全局不一致项;随后,采用凸距离目标函数估计初始相机位置和3D点,再通过非双线性角度目标函数优化。由于平移平均中旋转固定,若旋转不准将严重限制位置精度,因此进一步通过有界角度优化和重投影束调整,选择空间分布均衡的特征轨迹实现旋转与位置联合鲁棒优化。最后,使用所有可靠特征轨迹进行完整束调整,精修相机参数与3D点。在多种序列与无序真实数据集上的大量实验表明,本方法在精度、鲁棒性和可扩展性上均显著优于当前最优方法。

原文摘要 · Abstract (English)

Global Structure-from-Motion (SfM) offers advantages over incremental methods in terms of efficiency and error distribution. However, the task of translation averaging remains challenging. Many existing methods rely solely on relative translations or feature tracks, which either degrade under collinear camera motion or are susceptible to outliers. In this paper, we propose a novel hybrid explicit translation averaging framework that incorporates both relative translations and feature tracks. Specifically, we first refine the relative translations using global camera rotations and remove globally inconsistent relative translations. Next, we employ convex distance-based objective functions to estimate the initial camera positions and 3D points, followed by refinement using a non-bilinear angle-based objective function. Furthermore, since camera rotations are fixed during translation averaging, inaccurate camera rotations can severely limit the accuracy of camera positions. To address this issue, we then robustly refine both camera rotations and camera positions with selected feature tracks through bounded angle-based refinement and subsequent reprojection-based bundle adjustment. In this step, feature tracks are selected to maintain a balanced spatial distribution and improve optimization efficiency. Finally, we perform a complete bundle adjustment using all reliable feature tracks to refine the camera parameters and 3D points. Extensive experiments on various sequential and unordered real-world datasets demonstrate the superior accuracy, robustness, and scalability of our approach, outperforming state-of-the-art methods in both accuracy and computational efficiency.

三维重建结构光恢复相机位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。