arXiv:2603.07313cs.LGcs.AI2026-03

对抗性隐状态训练提升部分可观测强化学习的鲁棒性

Adversarial Latent-State Training for Robust Policies in Partially Observable Domains

  • 设计对抗性隐状态初始分布,模拟最坏情况下的环境变化
  • 在Battleship任务中,对抗训练使平均鲁棒性差距从10.3降至3.1次射击
  • 理论指导的诊断方法可识别优化过程中的实际限制与近似误差

部分可观测强化学习中的隐状态分布漂移仍具挑战性。本文提出对抗性隐状态初始状态POMDP框架,即对手在回合开始前选定隐藏的初始隐状态分布。理论上,我们证明了隐状态极小极大原理,刻画了最坏情况下的防御者分布,并推导出带有限样本集中界的确切最优响应不等式。在Battleship基准测试中,实验表明,对偏移隐状态分布的针对性暴露,将均匀分布与扩散分布间的平均鲁棒性差距从10.3次射击降低至3.1次,且在相同预算下实现。迭代最优响应训练表现出与理论预测一致的预算敏感行为,经考虑折扣PPO代理函数和有限样本噪声后,验证了其一致性。最终,该框架为隐状态初始分布问题提供了清晰的评估机制,并生成可解释的理论引导诊断,同时明确指出了实现层面代理函数与优化极限的影响所在。

原文摘要 · Abstract (English)

Robustness under latent distribution shift remains challenging in partially observable reinforcement learning. We formalize a focused setting where an adversary selects a hidden initial latent distribution before the episode, termed an adversarial latent-initial-state POMDP. Theoretically, we prove a latent minimax principle, characterize worst-case defender distributions, and derive approximate best-response inequalities with finite-sample concentration bounds that make the optimization and sampling terms explicit. Empirically, using a Battleship benchmark, we demonstrate that targeted exposure to shifted latent distributions reduces average robustness gaps between Spread and Uniform distributions from 10.3 to 3.1 shots at equal budget. Furthermore, iterative best-response training exhibits budget-sensitive behavior that is qualitatively consistent with the theorem-guided diagnostics once one accounts for discounted PPO surrogates and finite-sample noise. Ultimately, we show that for latent-initial-state problems, the framework yields a clean evaluation game and useful theorem-motivated diagnostics while also making clear where implementation-level surrogates and optimization limits enter.

强化学习鲁棒性对抗训练部分可观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。