arXiv:2601.08777cs.LGcs.AI2026-01

通过测试时扩展实现模型输出多样性,提升个性化对齐效果。

Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling

  • 引入多候选响应机制,利用测试时缩放增强对齐能力。
  • 理论证明最优收敛速率为 k/(k+1),超越现有方法。
  • 适合追求高可靠性与个性化生成的AI系统开发者。

将大语言模型对齐以满足用户异质且可能冲突的偏好,是实现个性化和可信AI的核心挑战。本文通过测试时缩放形式化理想化的通用对齐:每个提示生成 k≥1 个候选响应,用户选择偏好项。提出 (k,f(k))-鲁棒对齐,要求 k 输出模型在对抗单输出模型时胜率不低于 f(k),并定义渐进通用对齐(U-对齐),即当 k→∞ 时 f(k)→1。主要结果表明最优收敛速率为:存在一类单输出策略,其 k 采样联合策略可达到 U-对齐,速率 f(k)=k/(k+1),且任何方法无法普遍更快。我们发现主流后训练方法(如基于人类反馈的纳什学习,NLHF)根本性低估测试时缩放优势——尽管 NLHF 在 k=1 时最优,但其采样策略通常确定性,无法保证胜率超过 1/2,除非允许任意小的松弛。根源在于输出缺乏多样性:现有对齐方法易坍缩为单一主流偏好响应,使额外采样冗余。相反,我们的方法保持输出多样性,达到最优测试时缩放速率。特别地,我们提出一类对称多玩家对齐博弈,并证明任意 (k+1) 玩家博弈的对称纳什均衡策略均能实现最优 (k,k/(k+1)) 鲁棒对齐。最后,我们为自洽学习动态提供理论收敛保证,并将框架扩展至对手也生成多响应的情形。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI. We formalize an ideal notion of universal alignment through test-time scaling: for each prompt, the model produces $k\ge 1$ candidate responses and a user selects their preferred one. We introduce $(k,f(k))$-robust alignment, which requires the $k$-output model to have win rate $f(k)$ against any other single-output model, and asymptotic universal alignment (U-alignment), which requires $f(k)\to 1$ as $k\to\infty$. Our main result characterizes the optimal convergence rate: there exists a family of single-output policies whose $k$-sample product policies achieve U-alignment at rate $f(k)=\frac{k}{k+1}$, and no method can achieve a faster rate in general. We show that popular post-training methods, including Nash learning from human feedback (NLHF), can fundamentally underutilize the benefits of test-time scaling. Even though NLHF is optimal for $k=1$, sampling from the resulting (often deterministic) policy cannot guarantee win rates above $\tfrac{1}{2}$ except for an arbitrarily small slack. This stems from a lack of output diversity: existing alignment methods can collapse to a single majority-preferred response, making additional samples redundant. In contrast, our approach preserves output diversity and achieves the optimal test-time scaling rate. In particular, we propose a family of symmetric multi-player alignment games and prove that any symmetric Nash equilibrium policy of the $(k+1)$-player alignment game achieves the optimal $(k,\frac{k}{k+1})$-robust alignment. Finally, we provide theoretical convergence guarantees for self-play learning dynamics in these games and extend the framework to opponents that also generate multiple responses.

大模型对齐测试时缩放输出多样性博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。