比较肌电信号与视频,提升建筑机器人对工人动作意图的识别效率
Signals vs. Videos: Advancing Motion Intention Recognition for Human-Robot Collaboration in Construction
- 用肌电数据和视频分别训练模型,识别施工动作意图
- 视频模型准确率达94%,但耗时0.15秒;肌电模型87%准确率,仅需0.04秒
- 揭示两种数据的优劣,指导实际施工中智能部署
建筑领域人机协作(HRC)依赖机器人精准及时地识别工人动作意图,以保障安全并提升效率。现有研究缺乏对信号与视频两类数据模态在动作意图识别中的对比分析。为此,本研究采用深度学习方法,评估了表面肌电图(sEMG)与视频序列在干墙安装任务初期动作意图识别中的表现。基于卷积神经网络-长短期记忆网络(CNN-LSTM)的肌电模型实现约87%的准确率,平均预测时间仅为0.04秒;而使用预训练视频Swin Transformer结合迁移学习的视频模型,准确率达到94%,但平均预测时间为0.15秒。研究凸显了两种数据模态的独特优势与权衡,为实际施工项目中人机协作系统的系统化部署提供了依据。
原文摘要 · Abstract (English)
Human-robot collaboration (HRC) in the construction industry depends on precise and prompt recognition of human motion intentions and actions by robots to maximize safety and workflow efficiency. There is a research gap in comparing data modalities, specifically signals and videos, for motion intention recognition. To address this, the study leverages deep learning to assess two different modalities in recognizing workers' motion intention at the early stage of movement in drywall installation tasks. The Convolutional Neural Network - Long Short-Term Memory (CNN-LSTM) model utilizing surface electromyography (sEMG) data achieved an accuracy of around 87% with an average time of 0.04 seconds to perform prediction on a sample input. Meanwhile, the pre-trained Video Swin Transformer combined with transfer learning harnessed video sequences as input to recognize motion intention and attained an accuracy of 94% but with a longer average time of 0.15 seconds for a similar prediction. This study emphasizes the unique strengths and trade-offs of both data formats, directing their systematic deployments to enhance HRC in real-world construction projects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。