无惯性传感器的四旋翼仅用摄像头实现机身坐标系状态估计。
IMU-Free Body-Frame State Estimation with Sparse Scene Flow for Quadcopters

- 在机身坐标系中构建扩展卡尔曼滤波,利用静止特征点作隐式惯性参考。
- 通过多视角联合优化生成稀疏三维点云,含位置、速度及协方差信息。
- 点云密度自适应调节,聚焦关注区域,适合无外部基础设施的导航系统。
我们提出一种针对X型四旋翼的纯视觉状态估计算法,仅依赖标准双目相机对与电机推力指令,无需惯性传感器。系统在机身坐标系中运行,使用同步双目图像和电机推力信号,基于复合流形状态⟨SE(3), ℝ³, …⟩的连续-离散扩展卡尔曼滤波器,估计机身姿态、速度、角速度、重力与扰动。通过检测(FAST、Shi-Tomasi)、时序跟踪(SSD、Lucas-Kanade)与双目匹配(NCC)提取特征点,搜索区域由滤波器提供的姿态与点不确定性预测。通过卡方门控筛选出静止点输入滤波器。系统还通过四视图(两组双目在两个时间戳)全捆绑调整生成稀疏3D点云,联合估计位置与速度,以滤波器提供的相对姿态为先验。特征点不直接参与求解,其信息通过姿态先验体现。点云密度空间自适应:外部焦点点控制分配,关注区域密集覆盖,其余区域稀疏。输出为机身坐标系状态估计、校准的姿态变化与稀疏场景光流。该系统可作为下游世界模型的测量源,不依赖GPS、IMU或任何世界坐标系基础设施,但架构支持未来融合这些信息。
原文摘要 · Abstract (English)
We present a vision-only state estimation system for X-configuration quadcopters equipped with a canonical stereo camera pair and no inertial sensors. The system operates entirely in the body frame, requiring only synchronised stereo images and motor thrust commands. A continuous-discrete extended Kalman filter on a composite manifold state $\langle SE(3), \mathbb{R}^3, \ldots \rangle$ maintains estimates of body-frame pose, velocity, angular velocity, gravity, and disturbances, using stationary scene points as implicit inertial references. Feature points are detected (FAST, Shi-Tomasi), tracked temporally (SSD, Lucas-Kanade) and matched across cameras (NCC), with search regions predicted from filter-derived pose and point uncertainty. Chi-squared gating on the normalised innovation admits only stationary points to the filter. The system also produces a sparse 3D point cloud carrying per-point position, velocity and joint covariance. These come from a 4-view (two stereo pairs at two timestamps) full bundle adjustment that jointly estimates position and velocity from stereo disparity and temporal parallax, with the filter-derived relative pose as a prior. Feature points in the EKF do not enter the solver; their information is reflected through the pose prior. Point cloud density is spatially adaptive: an external focus point directs allocation, producing dense coverage in the region of attention and sparse coverage elsewhere. The output is a body-frame state estimate, a calibrated pose change, and a sparse scene flow. It is intended as a measurement source for a downstream world model anchored in the current body frame, without dependence on GPS, IMU, or any world-frame infrastructure, though the architecture accommodates their future integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。