arXiv:2604.03023cs.RO2026-04

让机器人控制既高性能又像真人驾驶,还能适应不同路况。

Behavior-Constrained Reinforcement Learning with Receding-Horizon Credit Assignment for High-Performance Control

  • 用滚动时域预测未来轨迹,动态调整奖励以约束行为偏差。
  • 在赛车模拟中跑出媲美职业车手的圈速,且驾驶风格高度一致。
  • 适合需要安全、稳定、类人决策的高动态控制场景。

在机器人控制中,学习高性能且与专家行为一致的策略是一项基本挑战。强化学习虽能发现高效策略,但常偏离理想的人类行为;而模仿学习受限于示范质量,难以超越专家数据。本文提出一种行为约束的强化学习框架,在提升性能的同时显式控制与专家行为的偏离。由于专家一致的行为本质上是轨迹级的,我们引入滚动时域预测机制,建模短期未来轨迹并在训练中提供前瞻奖励。为应对人类行为在扰动和变化条件下的自然差异,我们进一步将策略条件化于参考轨迹,使其可表示一组专家一致的行为分布,而非单一确定目标。我们在高保真赛车仿真中评估该方法,使用专业车手的数据,该领域以极端动态和狭窄性能裕度为特征。所学策略实现竞争性圈速,同时与专家驾驶行为高度对齐,优于基线方法在性能和模仿质量上的表现。此外,通过驾驶员在环仿真进行人因评估,结果表明所学策略再现了与顶级职业车手反馈一致的设置依赖型驾驶特征。这些结果证明,本方法可学习到既最优又行为一致的控制策略,能在复杂控制系统中可靠替代人类决策。

原文摘要 · Abstract (English)

Learning high-performance control policies that remain consistent with expert behavior is a fundamental challenge in robotics. Reinforcement learning can discover high-performing strategies but often departs from desirable human behavior, whereas imitation learning is limited by demonstration quality and struggles to improve beyond expert data. We propose a behavior-constrained reinforcement learning framework that improves beyond demonstrations while explicitly controlling deviation from expert behavior. Because expert-consistent behavior in dynamic control is inherently trajectory-level, we introduce a receding-horizon predictive mechanism that models short-term future trajectories and provides look-ahead rewards during training. To account for the natural variability of human behavior under disturbances and changing conditions, we further condition the policy on reference trajectories, allowing it to represent a distribution of expert-consistent behaviors rather than a single deterministic target. Empirically, we evaluate the approach in high-fidelity race car simulation using data from professional drivers, a domain characterized by extreme dynamics and narrow performance margins. The learned policies achieve competitive lap times while maintaining close alignment with expert driving behavior, outperforming baseline methods in both performance and imitation quality. Beyond standard benchmarks, we conduct human-grounded evaluation in a driver-in-the-loop simulator and show that the learned policies reproduce setup-dependent driving characteristics consistent with the feedback of top-class professional race drivers. These results demonstrate that our method enables learning high-performance control policies that are both optimal and behavior-consistent, and can serve as reliable surrogates for human decision-making in complex control systems.

强化学习控制策略行为一致性赛车仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。