arXiv:2511.12804cs.LGcs.AI2025-11AAAI被引 1

研究自演化模型中长期对齐的递归机制,揭示其三类收敛状态。

The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation

  • 用双阶段贝叶斯-特里模型构建递归筛选框架
  • 发现共识坍塌、共享最优妥协、不对称优化三类收敛态
  • 证明初始依赖不可消除,对齐是动态演化过程

在以自身输出为训练数据的自演化生成模型中,与用户偏好的对齐成为递归而非一次性过程。我们首次为这种递归重训练的长期影响提供形式化分析基础。基于基于贝叶斯-特里(Bradley-Terry, BT)模型的双阶段校准机制,我们将对齐建模为两方互动:模型所有者负责筛选哪些输出应被模型学习,公众用户则通过与模型交互决定哪些输出最终被分享和保留。分析揭示了三种结构收敛态,取决于偏好对齐程度:共识坍塌、共享最优妥协、不对称优化。我们证明了一个根本不可能性定理:任何基于递归BT的校准机制都无法同时保持多样性、确保对称影响力并消除对初始化的依赖。将该过程视为动态社会选择,表明对齐并非静态目标,而是受权力失衡与路径依赖共同塑造的演化均衡。

原文摘要 · Abstract (English)

In self-consuming generative models that train on their own outputs, alignment with user preferences becomes a recursive rather than one-time process. We provide the first formal foundation for analyzing the long-term effects of such recursive retraining on alignment. Under a two-stage curation mechanism based on the Bradley-Terry (BT) model, we model alignment as an interaction between two factions: the Model Owner, who filters which outputs should be learned by the model, and the Public User, who determines which outputs are ultimately shared and retained through interactions with the model. Our analysis reveals three structural convergence regimes depending on the degree of preference alignment: consensus collapse, compromise on shared optima, and asymmetric refinement. We prove a fundamental impossibility theorem: no recursive BT-based curation mechanism can simultaneously preserve diversity, ensure symmetric influence, and eliminate dependence on initialization. Framing the process as dynamic social choice, we show that alignment is not a static goal but an evolving equilibrium, shaped both by power asymmetries and path dependence.

对齐递归学习社会选择贝叶斯-特里

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。