arXiv:2603.01290cs.AIcs.GT2026-03

用隐马尔可夫模型推断对手电量状态,优化F1 2026赛季能量策略。

Opponent State Inference Under Partial Observability: An HMM-POMDP Framework for 2026 Formula 1 Energy Strategy

  • 构建两层框架:HMM推断对手电池状态与模式,DQN决策能量分配。
  • 对电池状态识别准确率达96.8%,对低耗能模式区分准确率89.4%。
  • 可发现对手的反收割陷阱,适合车队策略组和智能驾驶研究者。

2026年F1技术规则引入50/50动力分配机制,结合无限再生与驾驶员控制的覆盖模式,最优能量策略不仅依赖自身状态,还需推断对手隐藏状态。这构成一个部分可观测随机博弈,无法通过单智能体优化解决。本文提出可计算的两层推理与决策框架:第一层为40状态隐马尔可夫模型(HMM),基于六项公开遥测信号,推断对手能量回收系统(ERS)充电水平(四种模式:高、中、低再生、低降额)、覆盖模式状态及轮胎损耗状态;第二层为深度Q网络(DQN)策略,输入HMM信念状态并选择能量部署策略。本文形式化定义了‘反收割陷阱’——对手故意压制可见能耗信号以诱使对手误判,其检测需同时推理ERS水平与再生/降额子模式。在合成比赛中,该HMM实现96.8%的ERS水平识别准确率(随机基线25%),低再生与低降额模式分类准确率89.4%,反收割陷阱条件检测召回率达96.3%。季前分析表明赛道依赖的充电可用性(每圈1.0×至2.2×)为主要干扰因素,墨尔本赛道为最复杂验证环境。2026年澳大利亚大奖赛(3月8日)起将启动基于真实比赛遥测的Baum-Welch校准。

原文摘要 · Abstract (English)

The 2026 Formula 1 technical regulations introduce a fundamental change to energy strategy: under a 50/50 internal combustion engine / battery power split with unlimited regeneration and a driver-controlled Override Mode, the optimal energy deployment policy depends not only on a driver's own state but on the hidden state of rival cars. This creates a Partially Observable Stochastic Game that cannot be solved by single-agent optimisation methods. We present a tractable two-layer inference and decision framework. The first layer is a 40-state Hidden Markov Model (HMM) that infers a probability distribution over each rival's ERS charge level (four modes: H, M, L_harvest, L_derate), Override Mode status, and tyre degradation state from six publicly observable telemetry signals. The second layer is a Deep Q-Network (DQN) policy that takes the HMM belief state as input and selects between energy deployment strategies. We formally characterise the counter-harvest trap, a deceptive strategy in which a car deliberately suppresses observable deployment signals to induce a rival into a failed attack, and show that detecting it requires belief-state inference over both ERS level and the harvest/derate sub-mode. On synthetic races, the HMM achieves 96.8% ERS-level accuracy (random baseline 25%), classifies L_harvest vs. L_derate with 89.4% accuracy, and detects counter-harvest trap conditions with 96.3% recall. Pre-season analysis indicates circuit-dependent recharge availability (1.0x to 2.2x per lap) as the primary confound; Melbourne is the hardest-case validation environment. Baum-Welch calibration on 2026 race telemetry begins with the Australian Grand Prix (8 March 2026).

F1强化学习隐马尔可夫博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。