arXiv:2608.17030cs.ROcs.GR2026-08

用生理启发的最小奖励,让肌肉骨骼模型一小时内学会人类跑步。

Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation

论文配图:Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation
图 1 · 摘自论文原文
  • 以肌肉等长点为控制变量,通过伸展反射自动计算肌兴奋度。
  • 仅用一个简单奖励,在1小时内训练出类人奔跑动作。
  • 适合研究运动控制、生物物理仿真与强化学习融合的学者。

人体肌肉骨骼系统具有高度冗余性,导致基于强化学习的运动生成在高维动作空间中探索效率极低。为此,我们提出λ-保持控制器,受广泛支持的等长点(EP)假说启发,其控制变量为每块肌肉的等长点阈值长度λ,由拉伸反射规律自动计算肌肉兴奋度。在步态周期内保持λ值不变,显著降低策略查询频率。该方法首次使肌肉驱动的骨骼模型在单小时训练内,仅凭极简奖励实现类人冲刺。高效的探索机制不仅是一种工程优化,更根植于生理学,融合了等长点假说、间歇控制与最优反馈控制。本成果不仅在预测仿真中实现了类人行为建模,还推动了可学习的人类运动控制模型发展。

原文摘要 · Abstract (English)

The massive overactuation in the human musculoskeletal system makes it challenging to train musculoskeletal models to generate human-like motion via reinforcement learning, primarily because exploration in the resulting high-dimensional and redundant action space is extremely inefficient. To address this problem, we propose the $λ$-hold controller, inspired by the equilibrium-point (EP) hypothesis, which has been widely supported by extensive evidence from human motor control studies. The policy's control variable is the per-muscle EP threshold length $λ$, from which a stretch-reflex recruitment law computes the muscle excitations automatically. Holding each $λ$ over an interval of the gait phase also sharply reduces the frequency at which the policy must be queried. Consequently, the controller, to our knowledge for the first time, enables a muscle-actuated skeletal model to learn human-like sprinting using only a minimal reward within an hour of training. The efficient exploration through the proposed $λ$-hold controller is not merely an engineering trick but an approach grounded in physiology, bringing together the EP hypothesis, intermittent control, and optimal feedback control. Beyond encapsulating human-like behavior in predictive simulation, this achievement contributes to developing a learnable model of the human motor controller.

肌肉骨骼仿真强化学习运动控制生理启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。