arXiv:2605.22305cs.LG2026-05

解析经典障碍车问题,提出轻量高效的新控制策略

Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks

  • 基于切比雪夫多项式构建通用控制策略,从原理出发设计
  • 相比神经网络减少277倍参数,性能提升6.18倍,样本效率更高
  • 适合对实时性、可解释性要求高的低维控制场景

我们首次解析了强化学习中的经典基准问题——障碍车问题,推导出最优控制解,填补了36年来的空白。分析发现最优策略极为简单,但现代强化学习智能体与之仍存在显著差距。受此启发,我们提出切比雪夫策略(Chebyshev policies),一类从第一性原理出发的通用(稠密)策略类。该策略可作为神经网络的即插即用替代品,将遗憾(regret)降低6.18倍,同时仅需277倍更少的参数,显著提升样本效率、可解释性与实时性。在多个强化学习任务中验证,包括一个真实世界的非线性运动控制测试平台,其表现持续优于采用PPO、ARS和REINFORCE训练的神经网络。结果表明,切比雪夫策略为低维控制任务提供了一种高效且轻量的替代或补充方案。

原文摘要 · Abstract (English)

We analytically solve the Mountain Car problem, a canonical benchmark in RL, and derive an optimal control solution, closing a gap after 36 years. This enables us to reveal two surprising insights: The optimal control is quite simple, yet modern RL agents display a large gap to optimality. Motivated by the analysis of the optimal control, we introduce Chebyshev policies as a universal (i.e. dense) class of RL policies from first principles. They can be trained as drop-in replacements of neural nets, reducing the regret by a factor of 6.18, while requiring 277 times fewer parameters, fostering sample efficiency, explainability and realtime capability. Chebyshev policies are evaluated on further RL tasks, including a real-world nonlinear motion control testbed. They consistently improve performance over neural nets with PPO, ARS and REINFORCE. Our results demonstrate how Chebyshev policies offer a compelling and lightweight alternative or addition to neural nets for low-dimensional control tasks.

强化学习控制策略轻量化模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。