arXiv:2601.13100cs.LG2026-01被引 2

提出递归知识蒸馏的数学框架,证明其能稳定收敛到教师分布。

Recursive Meta-Distillation: An Axiomatic Framework for Iterative Knowledge Refinement

  • 用算子理论构建递归蒸馏的公理体系,确保多轮蒸馏有数学基础。
  • 在合理假设下,证明KL散度逐轮收缩,实现几何级数收敛。
  • 适用于多教师、多轮蒸馏场景,为稳定性分析提供通用理论工具。

概率域知识蒸馏的近期研究建立了温度缩放、多教师聚合与偏差-方差权衡的公理框架,但递归或多代蒸馏的数学行为仍不明确,以往方法多依赖经验启发。本文提出递归元蒸馏的公理化与算子理论框架,将迭代知识蒸馏建模为一系列带基教师锚定的概率分布算子。定义了有效元教师构建的结构公理,并证明存在非平凡算子族满足这些公理,无需指定具体算法或损失函数。在温和可实现性与凸性假设下,证明锚定的递归蒸馏诱导KL散度收缩,实现几何收敛至基教师分布,并存在唯一全局吸引不动点。该贡献为基础性而非算法性:框架刻画了递归蒸馏在何种条件下数学上良定且收敛,而非误差累积,独立于模型架构、优化细节或算子具体实现。结果为容量受限下迭代与多教师蒸馏的稳定性、偏差-方差行为及失败模式提供了理论依据。

原文摘要 · Abstract (English)

Recent work in probability-domain knowledge distillation has established axiomatic frameworks for temperature scaling, multi-teacher aggregation, and bias-variance trade-offs in single-stage settings. However, the mathematical behavior of recursive or multi-generation distillation remains poorly understood, with prior approaches relying primarily on empirical heuristics. In this work, we introduce an axiomatic and operator-theoretic framework for recursive meta-distillation, formalizing iterative knowledge distillation as a sequence of probability-distribution operators with explicit anchoring to base teachers. We define structural axioms for valid meta-teacher construction and prove the existence of non-trivial operator families satisfying these axioms without specifying particular algorithms or loss functions. Under mild realizability and convexity assumptions, we show that anchored recursive distillation induces contraction in KL divergence, yielding geometric convergence to base teacher distributions and a unique, globally attractive fixed point. The contribution is foundational rather than algorithmic: the framework characterizes when recursive distillation is mathematically well-posed and convergent rather than error-accumulating, independent of model architecture, optimization details, or specific operator instantiations. These results provide a theoretical basis for understanding stability, bias-variance behavior, and failure modes in iterative and multi-teacher distillation under capacity constraints.

知识蒸馏递归学习理论分析算子理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。