arXiv:2502.00620cs.LGcs.AI2025-02ICML被引 8

揭示弱模型指导强模型时的泛化机制,提出可预测性能的表征度量方法。

Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions

  • 用主成分分析的核函数刻画弱强模型表征差异,构建可学习空间。
  • 实验证明该方法在52个大模型上能准确预测弱到强泛化趋势。
  • 无需标签即可评估强模型潜力,适合大模型训练优化场景。

弱到强泛化(W2SG)指由弱模型监督更强模型,是理解人类未来引导超智能的重要类比。尽管已有实证表明强模型可超越其弱监督者,但其内在机制仍不清晰。本文从理论出发,提出通过弱、强模型内部表征的主成分生成核函数,定义一个空间,反映弱模型无法学习而强模型可掌握的内容。标签在此空间的投影量化了强模型因弱监督而未达潜力的程度。该理论揭示了某些弱监督错误可被强模型纠正,且不受过拟合影响。实验在分子预测与5项NLP任务中验证:基于表征的度量能有效预测52个大模型的W2SG表现,且无需标签。

原文摘要 · Abstract (English)

Weak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence in the future. Promising empirical results revealed that a strong model can surpass its weak supervisor. While recent work has offered theoretical insights into this phenomenon, a clear understanding of the interactions between weak and strong models that drive W2SG remains elusive. We investigate W2SG through a theoretical lens and show that it can be characterized using kernels derived from the principal components of weak and strong models' internal representations. These kernels can be used to define a space that, at a high level, captures what the weak model is unable to learn but is learnable by the strong model. The projection of labels onto this space quantifies how much the strong model falls short of its full potential due to weak supervision. This characterization also provides insights into how certain errors in weak supervision can be corrected by the strong model, regardless of overfitting. Our theory has significant practical implications, providing a representation-based metric that predicts W2SG performance trends without requiring labels, as shown in experiments on molecular predictions with transformers and 5 NLP tasks involving 52 LLMs.

模型泛化大模型表征分析监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。