arXiv:2603.07797cs.ROcs.LG2026-03

用统一成本函数解释人类伸手动作,比传统方法更准确高效。

Toward Global Intent Inference for Human Motion by Inverse Reinforcement Learning

  • 基于逆强化学习推断随时间变化的成本权重
  • 相比基线平均降低27%轨迹误差,提升预测精度
  • 适合研究运动控制、人机交互与生物力学建模者

本文探讨是否存在一个单一的统一成本函数,能够解释并预测人类伸手动作,而非依赖于个体或姿势特定的优化准则。采用最小观测逆强化学习(MO-IRL)算法,结合七维候选成本项,高效估计标准平面伸手任务中随时间变化的成本权重。相比双层优化方法,MO-IRL收敛速度提升数个数量级,且仅需少量数据即可探索时变成本结构。评估了三个层次的通用性:个体依赖姿态依赖、个体依赖姿态独立、个体独立姿态独立。在所有情况下,时变权重显著改善轨迹重建,平均相较基线减少27%的均方根误差(RMSE)。推断出的成本显示关节加速度调节起主导作用,辅以较小的力矩变化平滑贡献。总体表明,一个不依赖个体和姿态的时变成本函数可高精度预测人类伸手轨迹,支持此类运动存在统一最优性原理。

原文摘要 · Abstract (English)

This paper investigates whether a single, unified cost function can explain and predict human reaching movements, in contrast with existing approaches that rely on subject- or posture-specific optimization criteria. Using the Minimal Observation Inverse Reinforcement Learning (MO-IRL) algorithm, together with a seven-dimensional set of candidate cost terms, we efficiently estimate time-varying cost weights for a standard planar reaching task. MO-IRL provides orders-of-magnitude faster convergence than bilevel formulations, while using only a fraction of the available data, enabling the practical exploration of time-varying cost structures. Three levels of generality are evaluated: Subject-Dependent Posture-Dependent, Subject-Dependent Posture-Independent, and Subject-Independent Posture-Independent. Across all cases, time-varying weights substantially improve trajectory reconstruction, yielding an average 27% reduction in RMSE compared to the baseline. The inferred costs consistently highlight a dominant role for joint-acceleration regulation, complemented by smaller contributions from torque-change smoothness. Overall, a single subject- and posture-agnostic time-varying cost function is shown to predict human reaching trajectories with high accuracy, supporting the existence of a unified optimality principle governing this class of movements.

运动预测逆强化学习人体动作建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。