用可穿戴传感器与机器人摄像头同步手部动作,实现远距离多人交互中的指令源识别。
HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
- 融合机器人摄像头光学流与手戴惯性传感器信号,对齐时空特征
- 在34米距离下三人场景中达到92.32%准确率,超越前代方法48.44%
- 适用于真实机器人部署,支持自然手势的可靠指令识别
远距离人机交互(HRI)仍处于探索阶段,其中指令源识别(CSI)因多用户和距离导致的传感器模糊而极具挑战。本文提出HiSync,一种基于光学-惯性融合的框架,将手部运动作为绑定线索,通过对齐安装在机器人上的摄像头光学流与佩戴在手上的惯性测量单元(IMU)信号来实现。我们首先设计了一个由12名用户参与的自定义手势集,并在远距离多用户交互场景中收集了包含38个样本的多模态命令手势数据集。随后,HiSync从摄像头和IMU数据中提取频域手部运动特征,利用学习的CSINet对IMU信号去噪,实现模态间的时间对齐,并采用距离感知的多窗口融合策略计算细微自然手势的跨模态相似性,从而实现鲁棒的CSI。在最多34米、三人的场景中,HiSync达到92.32%的准确率,优于先前最先进方法48.44%。该系统也在真实机器人上进行了验证。通过实现可靠且自然的指令源识别,HiSync为公共空间中的交互提供了实用基础与设计指导。
原文摘要 · Abstract (English)
Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) - determining who issued a command - is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI. https://github.com/OctopusWen/HiSync
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。