用隐马尔可夫模型推断对手电量状态,优化F1 2026赛季能量策略。
Opponent State Inference Under Partial Observability: An HMM-POMDP Framework for 2026 Formula 1 Energy Strategy
- 构建两层框架:HMM推断对手电池状态与模式,DQN决策能量分配。
- 对电池状态识别准确率达96.8%,对低耗能模式区分准确率89.4%。
- 可发现对手的反收割陷阱,适合车队策略组和智能驾驶研究者。
2026年F1技术规则引入50/50动力分配机制,结合无限再生与驾驶员控制的覆盖模式,最优能量策略不仅依赖自身状态,还需推断对手隐藏状态。这构成一个部分可观测随机博弈,无法通过单智能体优化解决。本文提出可计算的两层推理与决策框架:第一层为40状态隐马尔可夫模型(HMM),基于六项公开遥测信号,推断对手能量回收系统(ERS)充电水平(四种模式:高、中、低再生、低降额)、覆盖模式状态及轮胎损耗状态;第二层为深度Q网络(DQN)策略,输入HMM信念状态并选择能量部署策略。本文形式化定义了‘反收割陷阱’——对手故意压制可见能耗信号以诱使对手误判,其检测需同时推理ERS水平与再生/降额子模式。在合成比赛中,该HMM实现96.8%的ERS水平识别准确率(随机基线25%),低再生与低降额模式分类准确率89.4%,反收割陷阱条件检测召回率达96.3%。季前分析表明赛道依赖的充电可用性(每圈1.0×至2.2×)为主要干扰因素,墨尔本赛道为最复杂验证环境。2026年澳大利亚大奖赛(3月8日)起将启动基于真实比赛遥测的Baum-Welch校准。
原文摘要 · Abstract (English)
The 2026 Formula 1 technical regulations introduce a fundamental change to energy strategy: under a 50/50 internal combustion engine / battery power split with unlimited regeneration and a driver-controlled Override Mode, the optimal energy deployment policy depends not only on a driver's own state but on the hidden state of rival cars. This creates a Partially Observable Stochastic Game that cannot be solved by single-agent optimisation methods. We present a tractable two-layer inference and decision framework. The first layer is a 40-state Hidden Markov Model (HMM) that infers a probability distribution over each rival's ERS charge level (four modes: H, M, L_harvest, L_derate), Override Mode status, and tyre degradation state from six publicly observable telemetry signals. The second layer is a Deep Q-Network (DQN) policy that takes the HMM belief state as input and selects between energy deployment strategies. We formally characterise the counter-harvest trap, a deceptive strategy in which a car deliberately suppresses observable deployment signals to induce a rival into a failed attack, and show that detecting it requires belief-state inference over both ERS level and the harvest/derate sub-mode. On synthetic races, the HMM achieves 96.8% ERS-level accuracy (random baseline 25%), classifies L_harvest vs. L_derate with 89.4% accuracy, and detects counter-harvest trap conditions with 96.3% recall. Pre-season analysis indicates circuit-dependent recharge availability (1.0x to 2.2x per lap) as the primary confound; Melbourne is the hardest-case validation environment. Baum-Welch calibration on 2026 race telemetry begins with the Australian Grand Prix (8 March 2026).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。