用强化学习优化F1赛车多智能体策略,实时响应对手动态。
Learning-based Multi-agent Race Strategies in Formula 1
- 基于自洽对抗训练和交互模块,学习能量、胎耗与进站决策。
- 能根据对手行为动态调整进站时机与轮胎选择,表现稳定可靠。
- 仅依赖真实比赛信息,适合赛前与赛中辅助决策使用。
在一级方程式赛车中,比赛策略需随赛况变化和对手行动不断调整。本文提出一种基于强化学习的多智能体策略优化方法。智能体学习在能量管理、轮胎损耗、气动相互作用及进站决策间取得平衡。基于预训练的单智能体策略,引入交互模块以建模对手行为。结合自洽对抗训练机制,生成具备竞争力的策略,智能体依据相对表现进行排名。实验结果表明,智能体能够根据对手动态调整进站时间、轮胎选择与能量分配,在多种情境下均展现出鲁棒且一致的赛车表现。由于该框架仅依赖真实比赛中可获取的信息,可有效支持赛前与赛中策略制定。
原文摘要 · Abstract (English)
In Formula 1, race strategies are adapted according to evolving race conditions and competitors' actions. This paper proposes a reinforcement learning approach for multi-agent race strategy optimization. Agents learn to balance energy management, tire degradation, aerodynamic interaction, and pit-stop decisions. Building on a pre-trained single-agent policy, we introduce an interaction module that accounts for the behavior of competitors. The combination of the interaction module and a self-play training scheme generates competitive policies, and agents are ranked based on their relative performance. Results show that the agents adapt pit timing, tire selection, and energy allocation in response to opponents, achieving robust and consistent race performance. Because the framework relies only on information available during real races, it can support race strategists' decisions before and during races.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。