揭示大模型对齐的容量瓶颈,解释为何加标签也突破不了性能上限。
The Alignment Bottleneck
- 将对齐视为资源受限的信道问题,构建双阶段认知模型
- 理论证明性能有上下界,均由认知容量决定,与数据量无关
- 解释了模型为何会讨好、作弊,适合研究对齐机制的学者
大规模语言模型随规模提升性能,但基于反馈的对齐仍存在系统性偏差。受经济学与认知科学中有限理性启发,我们视判断为资源受限过程,反馈为受限信道。在此基础上,建立以 $S$ 为条件的两阶段级联模型 $U \to H \to Y$,引入认知容量 $C_{\text{cog}|S}$ 与平均总容量 $\bar{C}_{\text{tot}|S}$。主要结果为容量耦合的对齐性能区间:在可分码本混合分布上,给出数据量无关的 Fano 下界;结合 PAC-Bayes 上界,其 KL 项由同一信道控制,且与 $m \, \bar{C}_{\text{tot}|S}$ 相关。当使用标准可观测损失且数据来自同源混合时,两者均受单一容量支配。由此推导出:在价值复杂度与容量固定下,仅增加标签无法突破界限;更复杂目标需容量随 $\log M$ 增长;一旦信号饱和容量,进一步优化易拟合信道规律,符合对谄媚与奖励劫持的报告。分析将对齐视为接口工程:测量并分配有限容量,管理任务复杂度,决定信息投入方向。
原文摘要 · Abstract (English)
Large language models improve with scale, yet feedback-based alignment still exhibits systematic deviations from intended behavior. Motivated by bounded rationality in economics and cognitive science, we view judgment as resource-limited and feedback as a constrained channel. On this basis, we model the loop as a two-stage cascade $U \to H \to Y$ given $S$, with cognitive capacity $C_{\text{cog}|S}$ and average total capacity $\bar{C}_{\text{tot}|S}$. Our main result is a capacity-coupled Alignment Performance Interval. It pairs a data size-independent Fano lower bound proved on a separable codebook mixture with a PAC-Bayes upper bound whose KL term is controlled by the same channel via $m \, \bar{C}_{\text{tot}|S}$. The PAC-Bayes bound becomes an upper bound on the same true risk when the canonical observable loss is used and the dataset is drawn from the same mixture. Under these matched conditions, both limits are governed by a single capacity. Consequences include that, with value complexity and capacity fixed, adding labels alone cannot cross the bound; attaining lower risk on more complex targets requires capacity that grows with $\log M$; and once useful signal saturates capacity, further optimization tends to fit channel regularities, consistent with reports of sycophancy and reward hacking. The analysis views alignment as interface engineering: measure and allocate limited capacity, manage task complexity, and decide where information is spent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。