arXiv:2602.23694cs.ROcs.AI2026-02

用传感器融合提升无人机与机器人手势操作的可靠性和可解释性

Interpretable Multimodal Gesture Recognition for Drone and Mobile Robot Teleoperation via Log-Likelihood Ratio Fusion

  • 结合手表惯性数据与手套电容信号,通过似然比融合提升识别准确率
  • 在20种手势上达到与视觉基线相当性能,计算成本降低70%以上
  • 适合需要实时、安全控制的灾后救援或工业巡检场景

人类操作员仍常需进入灾难现场和工业设施等危险环境,此时直观可靠的移动机器人与无人机远程操控至关重要。双手自由的操作方式可提升操作员机动性与态势感知,从而增强安全性。尽管基于视觉的手势识别已用于无手操控,但在遮挡、光照变化和复杂背景下性能显著下降,限制了其在真实任务中的应用。为此,本文提出一种多模态手势识别框架,融合双腕苹果手表的惯性数据(加速度计、陀螺仪、方向)与自研手套的电容传感信号。设计基于对数似然比(LLR)的晚期融合策略,不仅提升识别性能,还通过量化各模态贡献实现可解释性。为支持研究,构建了一个包含20种受飞机引导信号启发的手势的新数据集,包含同步的RGB视频、IMU和电容传感器数据。实验表明,该框架性能接近先进视觉基线,同时显著降低计算开销、模型规模与训练时间,适用于实时机器人控制。因此,强调了基于传感器的多模态融合在手势驱动的移动机器人与无人机遥控中的鲁棒性与可解释性潜力。

原文摘要 · Abstract (English)

Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and reliable teleoperation of mobile robots and Unmanned Aerial Vehicles (UAVs) is essential. In this context, hands-free teleoperation enhances operator mobility and situational awareness, thereby improving safety in hazardous environments. While vision-based gesture recognition has been explored as one method for hands-free teleoperation, its performance often deteriorates under occlusions, lighting variations, and cluttered backgrounds, limiting its applicability in real-world operations. To overcome these limitations, we propose a multimodal gesture recognition framework that integrates inertial data (accelerometer, gyroscope, and orientation) from Apple Watches on both wrists with capacitive sensing signals from custom gloves. We design a late fusion strategy based on the log-likelihood ratio (LLR), which not only enhances recognition performance but also provides interpretability by quantifying modality-specific contributions. To support this research, we introduce a new dataset of 20 distinct gestures inspired by aircraft marshalling signals, comprising synchronized RGB video, IMU, and capacitive sensor data. Experimental results demonstrate that our framework achieves performance comparable to a state-of-the-art vision-based baseline while significantly reducing computational cost, model size, and training time, making it well suited for real-time robot control. We therefore underscore the potential of sensor-based multimodal fusion as a robust and interpretable solution for gesture-driven mobile robot and drone teleoperation.

手势识别多模态融合无人机控制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。