arXiv:2608.15402cs.LG2026-08

提出推理时对齐的理论框架,解决未知奖励下的模型优化问题。

Towards a theory of inference-time alignment with unknown rewards

  • 将对齐建模为弱到强学习问题,无需预设奖励函数。
  • 引入对齐维度概念,其有限性决定对齐是否可学习。
  • 通过成对比较和锦标赛机制实现高效响应生成,适合理论研究者。

生成模型对齐受到广泛关注,监督微调和推理时计算已取得显著进展,但其统计学习视角仍不清晰。本文将推理时对齐形式化为弱到强学习问题:假设参考策略(弱模型)表现良好,目标是生成强模型,在测试时以任意高概率预测出优质响应。该问题从零开始学习,所有内容均来自数据,不依赖于良好奖励估计,与现有推理时对齐理论不同。我们的框架与Joshi等人(arXiv:2510.15464)的工作相似,即每个提示可能存在多个优质响应。对齐可学习性遵循标准PAC学习原则。我们引入奖励类的新组合维度——对齐维度,并证明其完全刻画对齐可学习性:奖励类对齐可学习当且仅当其对齐维度有限。核心学习过程通过学习成对比较器,再对候选响应进行锦标赛排序实现。我们认为这些结果可能为对齐的完整理论理解提供启示。

原文摘要 · Abstract (English)

Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning perspective. We formulate inference-time alignment as a weak-to-strong learning problem, where a reference policy (weak model) is assumed to be fairly good and the goal is to produce a strong model that predicts a good response at test time with arbitrarily high probability. Our problem is formulated as learning from scratch --- everything is learned from data rather than assuming access to a good reward estimate, and thus differs from the existing inference-time alignment theory. Our framework shares similarity to the recent work of Joshi et al., (arXiv:2510.15464), where for each prompt, there could be multiple good responses. Our definition of the alignment learnability follows the standard PAC learning principle. We introduce a novel combinatorial dimension of the reward class which we call the alignment dimension, and show that it completely characterizes the alignment learnability --- a reward class is alignment learnable if and only if its alignment dimension is finite. The core of our learning procedure works by learning a pairwise comparator and then running a tournament over candidate responses. We believe that our results might shed light toward establishing a complete theoretical understanding of alignment.

对齐理论强化学习机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。