解决语言智能体对战中策略僵局问题,让对话持续进化
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

- 通过双尺度价值基线与对局熵检测策略僵局
- 动态调整优势函数,恢复策略优化梯度
- 在多类社交对话游戏中实现持续进化
尽管基于可验证奖励的强化学习(RLVR)在封闭任务中表现良好,但在开放式的社交语言游戏中通过自对弈扩展时暴露出一个关键问题:进化僵局。由于策略空间庞大,语言智能体常收敛至同质化行为,导致对局结果确定性高,丧失策略演化的梯度信号。为此,我们提出双尺度演化策略训练(DEPT)。DEPT引入时间尺度演化感知机制,通过量化双尺度价值基线差异与对局熵来检测僵局;一旦感知到崩溃,便激活非对称优势重塑,动态调节优化景观以实现干预。该方法有效恢复梯度信号,维持持续的战略探索。在多个社交语言游戏上的大量实验表明,DEPT优于强基准模型,避免策略退化,驱动社交语言智能体的持续演化。
原文摘要 · Abstract (English)
While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for policy evolution. To tackle this issue, we propose Dual-scale Evolutionary Policy Training (DEPT) for social language games. DEPT introduces a time-scaled evolutionary perception mechanism that detects impasse by quantifying dual-scale value baseline divergence alongside match entropy. Upon perceiving the collapse, it then activates asymmetric advantage reshaping to dynamically modulate the optimization landscape for intervention. Thus, our method effectively restores gradient signals and enforces sustained strategic exploration. Extensive experiments on multiple social language games demonstrate that DEPT outperforms strong baselines, avoiding policy degeneration and driving the continuous evolution of social language agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。