融合触觉与运动数据,提升机器人协作中的动作识别精度
A Comparative Study of Human Activity Recognition: Motion, Tactile, and multi-modal Approaches
- 结合触觉与运动传感器,构建多模态识别框架
- 多模态方法在15类动作识别中准确率最高,优于单一模态
- 适合需要高精度动作理解的协作机器人场景
人体动作识别(HAR)对人机协作(HRC)至关重要,使机器人能够解析并响应人类行为。本研究评估了基于视觉的触觉传感器对15种动作的分类能力,并与基于惯性测量单元(IMU)的数据手套进行比较。提出一种融合触觉与运动数据的多模态框架,对比三种方法:基于运动数据的动作分类(MBC)、基于单/双视频流的触觉分类(TBC)以及融合两者数据的多模态分类(MMC)。通过分段数据集的离线验证和连续动作序列的在线验证,评估各配置在受控条件下的性能。结果表明,多模态方法始终优于单一模态,凸显了触觉与运动感知融合在协作机器人中提升动作识别系统的潜力。
原文摘要 · Abstract (English)
Human activity recognition (HAR) is essential for effective Human-Robot Collaboration (HRC), enabling robots to interpret and respond to human actions. This study evaluates the ability of a vision-based tactile sensor to classify 15 activities, comparing its performance to an IMU-based data glove. Additionally, we propose a multi-modal framework combining tactile and motion data to leverage their complementary strengths. We examined three approaches: motion-based classification (MBC) using IMU data, tactile-based classification (TBC) with single or dual video streams, and multi-modal classification (MMC) integrating both. Offline validation on segmented datasets assessed each configuration's accuracy under controlled conditions, while online validation on continuous action sequences tested online performance. Results showed the multi-modal approach consistently outperformed single-modality methods, highlighting the potential of integrating tactile and motion sensing to enhance HAR systems for collaborative robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。