arXiv:2608.13918cs.CV2026-08

无需控制点也能高精度估算相机平台相对运动,适合长距离视觉监测。

Beyond Control Points: Arcsecond Relative-Motion Estimation of Vision Measurement Platforms With Incomplete or Absent Control Fields

论文配图:Beyond Control Points: Arcsecond Relative-Motion Estimation of Vision Measurement Platforms With Incomplete or Absent Control Fields
图 1 · 摘自论文原文
  • 直接从图像位移和已知三维点估计平台运动,不依赖非线性优化或初始姿态。
  • 无控制点时旋转误差仅2.97弧秒,一个控制点时平移误差1.19毫米。
  • 对控制场污染免疫,计算快至0.46毫秒,适合实时高精度监测场景。

远距离基于视觉的形变监测对相机平台运动极为敏感。传统绝对姿态差分依赖专用控制数据,会引入两个独立的姿态误差。本文提出一种自适应控制的差分框架,直接从图像位移和已知3D点估计帧间平台运动。无专用控制点时,可从观测点恢复平台旋转;一个控制点可实现平移的先验约束恢复,两个非平行控制光线可完全恢复3D平移。该框架无需非线性优化且无需初始姿态估计。排除控制数据参与旋转阶段,使旋转估计完全免疫于控制场内污染。继承的差分形式可精确抵消平移外参误差。我们推导了旋转可观测性条件、未建模平移与非刚性点运动的泄漏界,以及单点轴向先验偏置律。在0.5像素图像噪声下,姿态变化达30~弧分,3D点扰动最大2毫米时,多相机估计算器的旋转均方根误差为2.97~弧秒,平均运行时间仅0.46毫秒。带一个控制点时,其先验约束平移均方根误差为1.19毫米。在无稳定控制场的桥梁实验中,相对于全站仪测量的坐标偏差均方根误差中位数为0.85毫米。该估计算器在公开的RGB-D和立体序列上,面对3D坐标扰动仍保持零发散。结果表明,其在精度、标定鲁棒性和计算效率方面均达到当前最优水平。

原文摘要 · Abstract (English)

Long-range vision-based deformation monitoring is highly sensitive to motion of the camera platform. Absolute-pose differencing typically relies on dedicated control data and propagates two independent pose errors into the relative-motion estimate. We develop a control-adaptive differential framework that estimates inter-frame platform motion directly from image displacements and known 3D points. With no dedicated control point, the framework recovers platform rotation from measurement-point observations. One surveyed control point enables prior-constrained translation recovery, while two nonparallel control rays recover full 3D translation. The framework requires neither nonlinear optimization nor an initial pose estimate. Excluding control data from the rotation stage makes the rotation estimate exactly immune to contamination confined to the control field. The inherited differential formulation also cancels translational extrinsic errors exactly. We derive the rotation observability condition, a leakage bound for unmodeled translation and nonrigid point motion, and the single-point axial-prior bias law. Under 0.5-pixel image noise, attitude changes of up to 30~arcmin, and 3D point perturbations of up to 2~mm, the multi-camera estimator achieves a rotation RMSE of 2.97~arcsec and an average runtime of 0.46~ms. With one surveyed control point, its prior-constrained translation RMSE is 1.19~mm. In a bridge experiment without a stable control field, the median coordinate-wise displacement RMSE relative to total-station measurements is 0.85~mm. The estimator also maintains zero divergence under the tested 3D coordinate perturbations on public RGB-D and stereo sequences. These results establish state-of-the-art accuracy, calibration robustness, and computational efficiency among the evaluated methods.

视觉测量相对运动姿态估计误差抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。