arXiv:2602.01357cs.LG2026-02被引 1

揭示自对弈训练本质是对抗式模仿学习,提升模型对齐稳定性。

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning

  • 将自对弈建模为模型与自身奖励玩家的博弈,统一理论框架。
  • 证明算法收敛至均衡点,且新方法在多个任务上性能更优。
  • 适合研究大模型对齐机制或优化自对弈训练的开发者参考。

自对弈后训练已成为微调大语言模型、将弱模型转化为强模型的有效方法,无需偏好数据。然而,其理论基础仍不清晰。本文通过将自对弈微调视为模型与由模型自身参数化的正则化隐式奖励玩家之间的极小极大博弈,将其与对抗式模仿学习相联系。该视角统一了自对弈模仿与通用偏好对齐。在此框架下,我们进行了博弈论分析,证明自对弈微调将收敛至均衡点。基于此理论,我们提出一种基于χ²散度变分目标、具有有界奖励的新自对弈模仿微调算法,显著提升稳定性。在多个语言模型微调任务上的实验表明,该方法持续优于现有自对弈方法,并验证了理论洞察。

原文摘要 · Abstract (English)

Self-play post-training methods has emerged as an effective approach for finetuning large language models and turn the weak language model into strong language model without preference data. However, the theoretical foundations for self-play finetuning remain underexplored. In this work, we tackle this by connecting self-play finetuning with adversarial imitation learning by formulating finetuning procedure as a min-max game between the model and a regularized implicit reward player parameterized by the model itself. This perspective unifies self-play imitation and general preference alignment within a common framework. Under this formulation, we present a game-theoretic analysis showing that the self-play finetuning will converge to it's equilibrium. Guided by this theoretical formulation, we propose a new self-play imitation finetuning algorithm based on the $χ^2$-divergence variational objective with bounded rewards and improved stability. Experiments on various of language model finetuning tasks demonstrate consistent improvements over existing self-play methods and validate our theoretical insights.

自对弈模仿学习大模型对齐博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。