用模糊推理从目标框几何特征实现可解释的无人机航向控制。
Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry

- 基于目标框的中心、面积和长宽比,用模糊系统生成航向指令。
- 在6169个样本上测试,误差低至0.14度,99.7%精度达±1度。
- 模型轻量可解释,适合资源受限的无人机实时导航。
基于视觉的无人机(UAV)对无人地面车辆(UGV)引导支持空地协作机器人,但受感知不确定性、计算资源有限及可解释性需求限制,机载视觉持续航向估计仍具挑战。现有深度学习与几何重建方法常需大量数据、外部定位或复杂建模,降低透明度与部署适用性。本文提出一种可解释的模糊推理框架,仅通过YOLO检测框提取的低维特征——目标中心位置、面积与长宽比——生成连续航向命令,无需显式几何建模。采用Mamdani模糊系统作为基准,后接一阶Takagi-Sugeno模型,每输入三个前提隶属函数,参数由训练集分位数推导,形成紧凑27规则结构。在6,169个带标签样本的VICON动作捕捉环境评估中,跨五次随机训练-测试划分,该模型测试集平均绝对误差为0.140° ± 0.003°,均方根误差0.200° ± 0.008°,最大绝对误差1.254° ± 0.121°。±1°阈值内准确率99.676% ± 0.270%,±3°与±5°阈值内准确率均为100.000% ± 0.000%。图像平面水平位移方向与预测航向符号的一致性达90.254% ± 0.612%。结果表明该框架透明、数据高效、计算轻量,适用于实时视觉引导无人机追踪移动地面目标。
原文摘要 · Abstract (English)
Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotics, but reliable continuous yaw estimation from onboard vision remains challenging because of sensing uncertainty, limited computation, and the need for interpretable control. Existing deep-learning and geometric-reconstruction approaches often require large datasets, external localization, or complex modeling assumptions, reducing transparency and deployment suitability on resource-constrained platforms. We present an interpretable fuzzy-inference framework that generates continuous yaw commands from low-dimensional features extracted from YOLO boxes: target centroid location, area, and aspect ratio. No explicit geometric modeling is required. A Mamdani fuzzy system serves as an interpretable baseline using a shoulder--triangle--shoulder input partition. It is followed by a first-order Takagi--Sugeno model with three antecedent membership terms per input, whose parameters are derived from training-set quantiles, yielding a compact 27-rule structure. Evaluation uses 6{,}169 labeled samples from a VICON motion-capture environment. Across five randomized train--test splits, the Takagi--Sugeno model achieves a test-set mean absolute error of $0.140^\circ \pm 0.003^\circ$, a root mean squared error of $0.200^\circ \pm 0.008^\circ$, and a maximum absolute error of $1.254^\circ \pm 0.121^\circ$. Within-threshold accuracies are $99.676% \pm 0.270%$ for $\pm1^\circ$ and $100.000% \pm 0.000%$ for both $\pm3^\circ$ and $\pm5^\circ$. Directional consistency between image-plane horizontal displacement and predicted yaw sign reaches $90.254% \pm 0.612%$. These results show that the framework is transparent, data-efficient, computationally lightweight, and suitable for real-time vision-based UAV guidance toward mobile ground targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。