arXiv:2508.17018stat.MLcs.LG2025-08被引 2

研究弱到强泛化中修正方法的局限性,发现其无法突破理论瓶颈。

Limitations of refinement methods for weak to strong generalization

  • 基于概率假设分析标签修正与弱训练的性能边界。
  • 两类方法均存在不可消除的误差,与理想模型有固定差距。
  • 为设计更优的弱到强泛化方法提供理论依据,适合相关研究者。

标准的大语言模型对齐技术依赖人工标注数据,可能限制对齐后模型达到人类水平的能力。标签修正和弱训练作为应对超对齐问题的有前景策略被提出。本文采用常用于研究标签修正的概率假设,分析修正方法是否可被其他方法超越,包括计算上不可行的理想模型(oracle)。结果表明,弱训练与标签修正均存在不可消除的误差,导致其性能与理想模型之间存在固定差距。这些发现推动未来研究探索融合标签修正或弱训练实用性与理想过程最优性的新型弱到强泛化方法。

原文摘要 · Abstract (English)

Standard techniques for aligning large language models (LLMs) utilize human-produced data, which could limit the capability of any aligned LLM to human level. Label refinement and weak training have emerged as promising strategies to address this superalignment problem. In this work, we adopt probabilistic assumptions commonly used to study label refinement and analyze whether refinement can be outperformed by alternative approaches, including computationally intractable oracle methods. We show that both weak training and label refinement suffer from irreducible error, leaving a performance gap between label refinement and the oracle. These results motivate future research into developing alternative methods for weak to strong generalization that synthesize the practicality of label refinement or weak training and the optimality of the oracle procedure.

大模型对齐弱监督理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。