arXiv:2606.01000cs.LGcs.CL2026-06

用可信度函数筛选弱教师标签,实现近乎无损的强学生训练

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher

论文配图:Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
图 1 · 摘自论文原文
  • 为每个弱标签分配可信度分数,仅用高可信标签训练
  • 在多个任务上表现接近甚至超过真实标签监督
  • 支持迭代链式训练,持续提升模型性能

弱教师到强学生泛化研究旨在利用稀缺可靠标签时,通过弱教师的监督提升强学生。本文将其视为数据选择问题,核心挑战在于识别哪些弱标签足够可信以作为训练信号。为此,提出可信度函数,为每个弱标签分配标量可信度分数,并据此过滤弱监督信号。在世界知识、定量推理和策略游戏等多个领域,可信度过滤使学生模型表现达到甚至超越真实标签监督水平,实现近乎无损的弱到强泛化。此外,可信度函数支持迭代弱到强链式训练,通过将学生重用于下一阶段教师,持续放大性能增益。可信度函数的优势可归因于多种机制。

原文摘要 · Abstract (English)

Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view this primarily as a data selection problem, where the key challenge is to identify which weak labels are reliable enough to serve as a training signal. To address this, we introduce trust functions that assign each weak label a scalar trust score and use these scores to filter weak supervision. Across several domains, including world knowledge, quantitative reasoning, and strategy games, trust filtering yields students that match and sometimes surpass ground-truth supervision, achieving near-lossless weak-to-strong generalization. Moreover, trust functions enable an iterative weak-to-strong chain that compounds gains by training a student and reusing it as the next teacher, amplifying the gains. There are several mechanisms to which advantage of trust functions can be attributed.

弱监督可信度泛化迭代训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。