让教师模型更真实地表达不确定,提升学生模型的泛化与可靠性。
Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty
- 用校准的不确定性引导教师输出,避免过度自信
- 在多种任务上使学生模型准确率更高、校准性更好
- 特别适合高类别数、长尾分布等复杂场景
知识蒸馏的核心在于传递教师模型蕴含的‘暗知识’——即类别间细微的概率关联和不确定性分布。然而,传统交叉熵训练的教师模型常产生尖锐且过度自信的输出,导致概率分布坍缩,难以传递有效信息,尤其在高类别任务中削弱了学生模型的表征学习能力。此外,这种不校准的输出在分布偏移下会降低鲁棒性。为此,我们提出校准不确定性蒸馏(CUD),从分布角度重构蒸馏过程,通过主动塑造教师预测分布,使其在应不确定处显露出可信的不确定性,从而为学生提供校准而非强化的指导信号。实验表明,该方法在多个基准上显著提升学生模型的准确率与校准性,尤其在分布偏移和长尾输入下表现更可靠。
原文摘要 · Abstract (English)
The core of knowledge distillation lies in transferring the teacher's rich 'dark knowledge'-subtle probabilistic patterns that reveal how classes are related and the distribution of uncertainties. While this idea is well established, teachers trained with conventional cross-entropy often fail to preserve such signals. Their distributions collapse into sharp, overconfident peaks that appear decisive but are in fact brittle, offering little beyond the hard label or subtly hindering representation-level transfer. This overconfidence is especially problematic in high-cardinality tasks, where the nuances among many plausible classes matter most for guiding a compact student. Moreover, such brittle targets reduce robustness under distribution shift, leaving students vulnerable to miscalibration in real-world conditions. To address this limitation, we revisit distillation from a distributional perspective and propose Calibrated Uncertainty Distillation (CUD), a framework designed to make dark knowledge more faithfully accessible. Instead of uncritically adopting the teacher's overconfidence, CUD encourages teachers to reveal uncertainty where it is informative and guides students to learn from targets that are calibrated rather than sharpened certainty. By directly shaping the teacher's predictive distribution before transfer, our approach balances accuracy and calibration, allowing students to benefit from both confident signals on easy cases and structured uncertainty on hard ones. Across diverse benchmarks, CUD yields students that are not only more accurate, but also more calibrated under shift and more reliable on ambiguous, long-tail inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。