arXiv:2607.01008eess.IVcs.RO2026-07

利用图像域俯仰翻滚信息增强无人机跟踪预测精度

Image-Domain Tilt Constrained Distributed Fusion for Maneuvering UAV Tracking with Multi-Camera Electro-Optical Observations

论文配图:Image-Domain Tilt Constrained Distributed Fusion for Maneuvering UAV Tracking with Multi-Camera Electro-Optical Observations
图 1 · 摘自论文原文
  • 从视频与传感器数据生成目标姿态标签,用YOLO检测实时获取位置和姿态
  • 引入图像域俯仰/翻滚作为加速度伪观测,使预测误差降低60.75%
  • 适用于多相机异步融合,抗抖动与系统偏差,适合复杂机动场景

短时预测对光电无人机跟踪至关重要,尤其在目标小、机动性强或间歇观测时。图像中心、视线及距离测量可直接约束位置,但对加速度约束较弱,导致剧烈机动时预测滞后。本文提出一种基于图像域俯仰翻滚约束的分布式融合方法。通过同步视频、云台IMU与无人机IMU数据,构建弱先验自动标注流程,生成带方向的边界框与图像域姿态标签。训练YOLO-OBB检测器实现在线目标位置与姿态估计。前端代码已开源。融合阶段采用位置-速度-加速度状态模型,将图像域俯仰与翻滚作为加速度相关伪观测。针对分布式跟踪,融合一台移动云台相机与两台固定地面相机,增广相机姿态误差状态以补偿外参漂移与跨相机系统不一致性。采用基于时间间隔的马氏距离门限与协方差扩展机制,有效剔除误检并处理丢失。仿真结果表明,加入俯仰/翻滚观测后,预测均方根误差由1.991米降至0.821米,累积误差减少60.75%。真实实验中自洽性评估显示累积误差下降18.10%。结果验证图像域姿态可为鲁棒短时预测提供有效加速度约束。

原文摘要 · Abstract (English)

Short-horizon prediction is essential for electro-optical UAV tracking, especially when the target is small, maneuvering, or intermittently observed. Image center, line-of-sight, and range measurements provide direct constraints on target position, but their constraints on acceleration are weak. As a result, prediction can lag during aggressive maneuvers. This paper proposes an image-domain tilt constrained distributed fusion method for maneuvering UAV tracking. The method uses the apparent roll and pitch of a rotorcraft target in the image as low-level maneuver cues. A weak-prior auto-labeling pipeline first generates oriented bounding box and image-domain tilt labels from synchronized video, gimbal IMU, and UAV IMU data. A YOLO-OBB detector is then trained to provide online target position and tilt measurements. The front-end Python implementation is publicly available at github.com/ShineMinxing/PythonYOLO. In the fusion stage, the UAV state is modeled by position, velocity, and acceleration. Image-domain roll and pitch are introduced as acceleration-related pseudo-observations. For distributed tracking, one mobile gimbal camera and two fixed ground cameras are fused asynchronously. Camera attitude error states are augmented into the filter to absorb extrinsic drift and cross-camera systematic inconsistency. A Mahalanobis gate with time-since-last-valid covariance widening is used to reject false detections and handle dropouts. In simulation, adding roll/pitch observations reduces the prediction RMSE from 1.991 m to 0.821 m and decreases the cumulative prediction error by 60.75\%. In real distributed experiments, a self-consistency evaluation shows an 18.10\% reduction in cumulative prediction error. The results show that image-domain tilt can provide useful acceleration constraints for robust short-horizon UAV prediction.

无人机跟踪多源融合姿态估计卡尔曼滤波

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。