arXiv:2501.04068cs.LGcs.AI2025-01被引 6

用强化学习优化F1赛车策略,提升比赛排名并增强决策可信度。

Explainable Reinforcement Learning for Formula One Race Strategy

  • 基于强化学习构建赛车策略模型,替代传统人工规则与蒙特卡洛方法。
  • 在巴林2023年大奖赛模拟中平均排名达第5.33位,优于最优基线的第5.63位。
  • 通过特征重要性与反事实分析提升模型可解释性,适合车队策略团队使用。

一级方程式赛车中,车队需通过制定赛车策略来提升最终排名,而无法在比赛中更改车辆配置。本文提出一种强化学习模型RSRL(Race Strategy Reinforcement Learning),用于控制赛车策略的仿真,作为比传统硬编码规则和蒙特卡洛方法更快的替代方案。在模拟中,赛车表现相当于预期排名为第5.5位(P5.5,P1为第一,P20为最后),实际平均排名达到第5.33位,优于最佳基线的第5.63位。我们还通过训练实现了对单一赛道或多赛道性能的优先优化。此外,结合特征重要性、基于决策树的代理模型及反事实分析,增强用户对模型的信任。最后,通过实例展示该方法在真实场景中的应用价值,建立仿真与现实之间的联系。

原文摘要 · Abstract (English)

In Formula One, teams compete to develop their cars and achieve the highest possible finishing position in each race. During a race, however, teams are unable to alter the car, so they must improve their cars' finishing positions via race strategy, i.e. optimising their selection of which tyre compounds to put on the car and when to do so. In this work, we introduce a reinforcement learning model, RSRL (Race Strategy Reinforcement Learning), to control race strategies in simulations, offering a faster alternative to the industry standard of hard-coded and Monte Carlo-based race strategies. Controlling cars with a pace equating to an expected finishing position of P5.5 (where P1 represents first place and P20 is last place), RSRL achieves an average finishing position of P5.33 on our test race, the 2023 Bahrain Grand Prix, outperforming the best baseline of P5.63. We then demonstrate, in a generalisability study, how performance for one track or multiple tracks can be prioritised via training. Further, we supplement model predictions with feature importance, decision tree-based surrogate models, and decision tree counterfactuals towards improving user trust in the model. Finally, we provide illustrations which exemplify our approach in real-world situations, drawing parallels between simulations and reality.

强化学习F1赛车可解释性策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。