arXiv:2510.07884cs.CLcs.AI2025-10被引 2

用对比解码提升弱模型训练强模型的泛化能力。

Contrastive Weak-to-strong Generalization

  • 通过对比解码在对齐前后弱模型间生成更高质量样本
  • 在多个模型族上实现一致性能提升,显著减少噪声与偏差
  • 适合追求高效模型扩展与无监督训练的AI研究者

弱到强泛化为规模化大语言模型提供了一种有前景的方法:在无需人类反馈或显式奖励建模的前提下,利用对齐后的弱模型生成样本训练更强模型。然而,弱模型输出中的噪声和偏差限制了其鲁棒性与泛化能力。本文提出隐式奖励机制,通过对数似然比近似显式奖励,并揭示其与对比解码(Contrastive Decoding, CD)的结构等价性。基于此,我们提出对比弱到强泛化(ConG)框架,利用对齐前后的弱模型进行对比解码以生成高质量样本。该方法有效实现能力迁移、去噪与鲁棒性增强,显著缓解传统方法的局限性。在多种模型家族上的实验结果表明,ConG具有稳定且一致的性能提升,验证了其通用性与有效性。整体表明,ConG有望推动弱到强泛化发展,为迈向通用人工智能提供新路径。

原文摘要 · Abstract (English)

Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling. However, its robustness and generalization are hindered by the noise and biases in weak-model outputs, which limit its applicability in practice. To address this challenge, we leverage implicit rewards, which approximate explicit rewards through log-likelihood ratios, and reveal their structural equivalence with Contrastive Decoding (CD), a decoding strategy shown to reduce noise in LLM generation. Building on this connection, we propose Contrastive Weak-to-Strong Generalization (ConG), a framework that employs contrastive decoding between pre- and post-alignment weak models to generate higher-quality samples. This approach enables more reliable capability transfer, denoising, and improved robustness, substantially mitigating the limitations of traditional weak-to-strong methods. Empirical results across different model families confirm consistent improvements, demonstrating the generality and effectiveness of ConG. Taken together, our findings highlight the potential of ConG to advance weak-to-strong generalization and provide a promising pathway toward AGI.

模型泛化对比解码无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。