用深度学习预测搬重物时全身姿态,精度提升超20%。
Evaluating the Performance of Deep Learning Models in Whole-body Dynamic 3D Posture Prediction During Load-reaching Activities
- 用双向LSTM和Transformer建模动态搬重物时的全身姿态变化。
- 引入新损失函数使手臂与腿部预测误差分别降低8%和21%。
- Transformer模型长期预测误差仅41.4毫米,比LSTM高58%精度。
本研究探索深度神经网络在动态搬重物任务中全身姿态预测的应用。采用双向长短期记忆网络(BLSTM)和Transformer两种时间序列模型,基于20名正常体重健康男性在不同负载位置下完成204次搬重物任务的3D全身体态数据进行训练。输入包括手-负载位置、举重方式(俯身、全蹲、半蹲)、操作方式(单手/双手)、身高体重及任务前25%时段的体态坐标,用于预测剩余75%时段的体态。为提升预测精度,提出一种新方法:通过优化新代价函数强制保持各肢体段长度恒定。结果表明,该方法使臂部与腿部预测误差分别下降约8%和21%;采用Transformer架构的模型表现出更优的长期性能,其均方根误差为41.4毫米,比基于BLSTM的模型精度高出约58%。本研究证明了利用神经网络捕捉3D运动帧间时序依赖性的有效性,为理解与预测人工搬运过程中的运动动力学提供了新途径。
原文摘要 · Abstract (English)
This study aimed to explore the application of deep neural networks for whole-body human posture prediction during dynamic load-reaching activities. Two time-series models were trained using bidirectional long short-term memory (BLSTM) and transformer architectures. The dataset consisted of 3D full-body plug-in gait dynamic coordinates from 20 normal-weight healthy male individuals each performing 204 load-reaching tasks from different load positions while adapting various lifting and handling techniques. The model inputs consisted of the 3D position of the hand-load position, lifting (stoop, full-squat and semi-squat) and handling (one- and two-handed) techniques, body weight and height, and the 3D coordinate data of the body posture from the first 25% of the task duration. These inputs were used by the models to predict body coordinates during the remaining 75% of the task period. Moreover, a novel method was proposed to improve the accuracy of the previous and present posture prediction networks by enforcing constant body segment lengths through the optimization of a new cost function. The results indicated that the new cost function decreased the prediction error of the models by approximately 8% and 21% for the arm and leg models, respectively. We indicated that utilizing the transformer architecture, with a root-mean-square-error of 41.4 mm, exhibited approximately 58% more accurate long-term performance than the BLSTM-based model. This study merits the use of neural networks that capture time series dependencies in 3D motion frames, providing a unique approach for understanding and predict motion dynamics during manual material handling activities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。