arXiv:2510.11877cs.LGcs.GT2025-10中稿 · NeurIPS

用序列建模提升对抗性博弈中的鲁棒性,首次实现决策变换器的抗攻击优化。

Robust Adversarial Reinforcement Learning in Stochastic Games via Sequence Modeling

  • 基于阶段博弈建模,用纳什Q值引导Transformer策略生成
  • 在多种对抗随机博弈中显著提升最坏情况回报表现
  • 适合关注鲁棒强化学习与博弈对抗场景的研究者

Transformer作为一种强大的序列建模架构,已被用于解决顺序决策问题,尤其通过决策变换器(DT)实现以期望回报为条件的策略学习。然而,基于序列建模的强化学习方法在对抗环境下的鲁棒性尚未得到充分研究。本文提出保守对抗鲁棒决策变换器(CART),据我们所知是首个专为提升DT在对抗随机博弈中鲁棒性的框架。我们将每阶段主角与对手的互动建模为阶段博弈,其收益定义为后续状态上的期望最大值,从而显式纳入随机状态转移。通过将Transformer策略条件化于该阶段博弈的纳什Q值,CART生成的策略同时具备更低可被利用性(对抗鲁棒)和对转移不确定性更强的保守性。实验表明,CART实现了更准确的最小最大值估计,并在多种对抗随机博弈中持续获得更优的最坏情况回报。

原文摘要 · Abstract (English)

The Transformer, a highly expressive architecture for sequence modeling, has recently been adapted to solve sequential decision-making, most notably through the Decision Transformer (DT), which learns policies by conditioning on desired returns. Yet, the adversarial robustness of reinforcement learning methods based on sequence modeling remains largely unexplored. Here we introduce the Conservative Adversarially Robust Decision Transformer (CART), to our knowledge the first framework designed to enhance the robustness of DT in adversarial stochastic games. We formulate the interaction between the protagonist and the adversary at each stage as a stage game, where the payoff is defined as the expected maximum value over subsequent states, thereby explicitly incorporating stochastic state transitions. By conditioning Transformer policies on the NashQ value derived from these stage games, CART generates policy that are simultaneously less exploitable (adversarially robust) and conservative to transition uncertainty. Empirically, CART achieves more accurate minimax value estimation and consistently attains superior worst-case returns across a range of adversarial stochastic games.

强化学习对抗鲁棒序列建模博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。