arXiv:2510.00329cs.RO2025-10

用极少数据学习人类伸手动作的动态成本规律,提升预测精度与泛化能力。

Learning Human Reaching Optimality Principles from Minimal Observation Inverse Reinforcement Learning

  • 基于最小观测逆强化学习,分阶段学习七种成本函数的动态权重组合。
  • 仅需每姿态10次示范,关节角度误差低至5.6度,优于静态权重的10.4度。
  • 可跨被试通用,适合用于机器人运动控制与人机协同系统设计。

本文研究将最小观测逆强化学习(MO-IRL)应用于建模与预测具有时变成本权重的人体手臂伸展动作。基于平面双连杆生物力学模型及受试者执行指指点任务的高分辨率运动捕捉数据,将每条轨迹分割为多个阶段,并学习各阶段中七种候选成本函数的特定组合。MO-IRL在最大熵逆强化学习框架下通过缩放观测与生成轨迹迭代优化成本权重,显著减少所需演示次数与收敛时间。在每姿态10次示范下,六段和八段权重划分的平均关节角均方根误差(RMSE)分别为6.4°和5.6°,优于单一定态权重的10.4°。交叉验证剩余试验及首次实现的跨被试验证(对未见被试20次试验),预测误差均约为8°,表明模型具备良好泛化能力。学习到的权重在运动起止阶段强调最小化关节加速度,符合生物运动中的平滑性原则。结果表明,MO-IRL能高效揭示人类运动控制中动态、普适的成本结构,适用于类人机器人系统。

原文摘要 · Abstract (English)

This paper investigates the application of Minimal Observation Inverse Reinforcement Learning (MO-IRL) to model and predict human arm-reaching movements with time-varying cost weights. Using a planar two-link biomechanical model and high-resolution motion-capture data from subjects performing a pointing task, we segment each trajectory into multiple phases and learn phase-specific combinations of seven candidate cost functions. MO-IRL iteratively refines cost weights by scaling observed and generated trajectories in the maximum entropy IRL formulation, greatly reducing the number of required demonstrations and convergence time compared to classical IRL approaches. Training on ten trials per posture yields average joint-angle Root Mean Squared Errors (RMSE) of 6.4 deg and 5.6 deg for six- and eight-segment weight divisions, respectively, versus 10.4 deg using a single static weight. Cross-validation on remaining trials and, for the first time, inter-subject validation on an unseen subject's 20 trials, demonstrates comparable predictive accuracy, around 8 deg RMSE, indicating robust generalization. Learned weights emphasize joint acceleration minimization during movement onset and termination, aligning with smoothness principles observed in biological motion. These results suggest that MO-IRL can efficiently uncover dynamic, subject-independent cost structures underlying human motor control, with potential applications for humanoid robots.

逆强化学习运动建模人机协同动态成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。