arXiv:2601.09165cs.LG2026-01被引 3

提出多教师知识蒸馏的数学框架,统一理论基础并证明其有效性。

Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation

  • 基于五个公理构建概率域知识聚合的数学框架,不依赖特定公式。
  • 证明多教师聚合可降低随机方差与系统偏差,且在异质教师下仍有效。
  • 适用于多种前沿模型,为实际实现提供理论支持,适合研究者参考。

在稀疏知识蒸馏(Sparse-KD)的概率域蒸馏框架基础上,我们构建了一个公理化、算子理论驱动的多教师集成知识蒸馏框架。该框架不指定具体的聚合公式,而是定义了五条核心公理:凸性、正性、连续性、权重单调性与温度一致性,以约束有效的知识聚合算子。我们证明了满足这些公理的算子族存在且不唯一,表明多种不同机制均可遵循同一基础原则。在此框架下,我们建立了无需依赖具体算子的理论保证:当教师模型异质时,多教师聚合能同时减小随机方差与系统性监督偏差,并提供詹森型界、对数损失保证及安全衰减性质。对于在教师权重上线性聚合的算子,我们在标准独立性假设下进一步推导出经典集成方差缩减结果,并拓展至误差相关情形。该框架为来自多样化前沿模型的多教师蒸馏提供了坚实的理论基础,同时兼容多种有效实现策略。

原文摘要 · Abstract (English)

Building on the probability-domain distillation framework of Sparse-KD, we develop an axiomatic, operator-theoretic framework for multi-teacher ensemble knowledge distillation. Rather than prescribing a specific aggregation formula, we define five core axioms governing valid knowledge aggregation operators, encompassing convexity, positivity, continuity, weight monotonicity, and temperature coherence. We prove the existence and non-uniqueness of operator families satisfying these axioms, establishing that multiple distinct aggregation mechanisms conform to the same foundational principles. Within this framework, we establish operator-agnostic guarantees showing that multi-teacher aggregation reduces both stochastic variance and systematic supervisory bias under heterogeneous teachers, while providing Jensen-type bounds, log-loss guarantees, and safety attenuation properties. For aggregation operators linear in teacher weights, we further establish classical ensemble variance-reduction results under standard independence assumptions, with extensions to correlated-error regimes. The framework provides theoretical grounding for multi-teacher distillation from diverse frontier models while admitting multiple valid implementation strategies.

知识蒸馏多教师理论框架概率聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。