arXiv:2506.03109cs.LG2025-06被引 4

用f散度提升弱监督到强模型的泛化能力,效果更好且更省资源。

On Weak-to-Strong Generalization and f-Divergence

  • 引入f散度作为信息论框架,统一弱-强泛化损失
  • 理论证明不同f散度在泛化上等价,且具样本复杂度优势
  • 实测比传统方法更抗噪声,适合资源受限场景

弱-强泛化(W2SG)作为一种新范式,通过弱监督者指导强预训练模型的能力提升。现有方法常需额外弱模型或复杂流程,带来显著计算与内存开销。受f散度在机器学习中有效性启发,本文将其引入W2SG作为信息论损失框架。理论分析揭示不同f散度损失在W2SG中的根本局限性与等价性,结合样本复杂度边界与信息论洞察。实验表明,该框架可有效提升强模型的泛化性能与抗噪能力,且通用性强于常用KL散度。

原文摘要 · Abstract (English)

Weak-to-strong generalization (W2SG) has emerged as a promising paradigm for stimulating the capabilities of strong pre-trained models by leveraging supervision from weaker supervisors. To improve the performance of the strong model, existing methods often require additional weak models or complex procedures, leading to substantial computational and memory overhead. Motivated by the effectiveness of $f$-divergence loss in various machine learning domains, we introduce $f$-divergence as an information-theoretic loss function framework in W2SG. Our theoretical analysis reveals fundamental limitations and equivalence of different $f$-divergence losses in W2SG, supported by sample complexity bounds and information-theoretic insights. We empirically demonstrate that $f$-divergence loss, which generalizes widely-used metrics like KL divergence, effectively improves generalization and noise tolerance of the strong model in practice.

泛化能力f散度弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。