arXiv:2605.30854cs.MAcs.AI2026-05

让语言模型在博弈中更安全,避免被欺负或串通

Safe Equilibrium Policy Optimization for Strategic Agent Policies

论文配图:Safe Equilibrium Policy Optimization for Strategic Agent Policies
图 1 · 摘自论文原文
  • 用惩罚项显式约束可被利用、串通和外部化成本的风险
  • 在扑克博弈中实现零被利用优势,四领域提升安全性
  • 适合研究多智能体安全、博弈行为控制的学者使用

用强化学习微调的语言模型通常只优化任务奖励,忽视多智能体战略结构。由于这些智能体基于自然语言描述状态并自由生成动作,策略性失败模式——如利用弱对手、协调达成有害均衡、转嫁成本——与语言接口本身密不可分。我们提出安全均衡策略优化( exttt{SEPO}),在期望收益基础上加入对可利用性、串通风险和外部成本的显式惩罚。将 exttt{SEPO} 作为奖励信号用于组相对策略优化(GRPO),应用于经监督微调(SFT)后的 Gemma~4 E4B-it 与 Qwen~3.5-4B 模型。在五类战略场景下评估:重复囚徒困境、重复拍卖、两种谈判变体及克努扑克。 exttt{SEPO} 在克努扑克中使两模型均实现零被利用优势,在四个领域优于基线模型的安全性表现,并纠正了 SFT 引入的过度合作行为。在谈判中, exttt{SEPO} 达到正向安全结果,且所有谈判配置的归一化相对优势均为正值。消融实验表明,每轮计算可利用性是必要的:共享常数惩罚在 GRPO 优势归一化中抵消(常数控制变量性质),导致梯度为零。

原文摘要 · Abstract (English)

Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these agents condition on natural language game-state descriptions and emit actions through free-form generation, strategic failure modes -- exploiting weaker opponents, coordinating on harmful equilibria, and externalizing costs are inseparable from the language interface itself. We propose Safe Equilibrium Policy Optimization (\sepo{}), a training objective that augments expected payoff with explicit penalties for exploitability, collusion risk, and externality cost. We implement \sepo{} as a reward signal for Group Relative Policy Optimization (GRPO), applied to Gemma~4 E4B-it and Qwen~3.5-4B after supervised fine-tuning (SFT). Evaluated across five strategic domains: Iterated Prisoner's Dilemma, repeated auctions, two negotiation variants, and Kuhn Poker. \sepo{} achieves zero exploit-pool advantage in Kuhn Poker for both models, outperforms the base model on safety in four domains, and corrects the over-cooperative behavior introduced by SFT. In negotiation, \sepo{} achieves a positive-safety outcome and only the positive normalized relative advantage of any negotiation configuration. Ablation experiments confirm that per-rollout exploit computation is necessary: a shared constant penalty cancels in GRPO advantage normalization (constant control-variate property), producing zero gradient. To support further research in strategic safety for agents, we release our \href{https://anonymous.4open.science/r/sepo-2668/README.md}{code} and SFT datasets.

多智能体博弈安全语言模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。