arXiv:2410.04638cs.LGstat.ML2024-10ICLR被引 20

弱教师用错误标签教强学生,也能让其正确分类。

Provable Weak-to-Strong Generalization via Benign Overfitting

  • 用伪标签反向训练强模型,理论证明可实现有效泛化
  • 在高维数据下,模型要么准确分类,要么随机猜测
  • 揭示了用对数概率做弱监督的潜在优势

机器学习中的经典师生模型是强教师指导弱学生。本文研究相反情形:弱教师用不完美伪标签指导强学生,即弱到强泛化。我们在一个过参数化的斯皮克协方差模型中进行理论分析,假设弱教师的伪标签近似随机猜测。在高维高斯特征下,我们严格证明强学生在弱监督后存在两种渐近行为:(1)成功泛化;(2)表现如随机猜测。该方法有望推广至多标签分类。为此,我们建立了相关高斯变量最大值的紧下尾不等式,或具独立意义。结果表明,当可用原始对数概率时,使用它们进行弱监督更具价值。

原文摘要 · Abstract (English)

The classic teacher-student model in machine learning posits that a strong teacher supervises a weak student to improve the student's capabilities. We instead consider the inverted situation, where a weak teacher supervises a strong student with imperfect pseudolabels. This paradigm was recently brought forth by Burns et al.'23 and termed \emph{weak-to-strong generalization}. We theoretically investigate weak-to-strong generalization for binary and multilabel classification in a stylized overparameterized spiked covariance model with Gaussian covariates where the weak teacher's pseudolabels are asymptotically like random guessing. Under these assumptions, we provably identify two asymptotic phases of the strong student's generalization after weak supervision: (1) successful generalization and (2) random guessing. Our techniques should eventually extend to weak-to-strong multiclass classification. Towards doing so, we prove a tight lower tail inequality for the maximum of correlated Gaussians, which may be of independent interest. Understanding the multilabel setting reinforces the value of using logits for weak supervision when they are available.

泛化理论弱监督过参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。