提出多教师知识蒸馏自适应加权的公理化框架,提升模型鲁棒性与安全性。
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
- 构建跨粒度(令牌/任务/上下文)自适应加权的公理体系
- 证明加权算子存在且不唯一,支持梯度优化收敛与稳定性
- 适用于异构数据、分布偏移及安全约束场景,适合研究者参考
多教师知识蒸馏在提升模型鲁棒性、效率和安全性方面日益重要,但现有方法多依赖启发式或实现特定的加权策略。本文提出一种无需依赖具体算子的公理化框架,覆盖令牌、任务和上下文三个互补尺度的自适应加权。我们形式化了加权算子良定义的结构条件,证明其存在性与非唯一性,并可基于乘积结构归一化进行分层组合。在该框架下,我们建立了标准假设下的梯度优化收敛性,分析了稳定性与扰动鲁棒性,并给出安全约束蒸馏的抽象表述。结果将理论保证与具体加权公式解耦,为异构环境、分布偏移及安全约束下的自适应蒸馏提供原则性分析基础。
原文摘要 · Abstract (English)
Knowledge distillation with multiple teachers is increasingly used to improve robustness, efficiency, and safety, yet existing approaches rely largely on heuristic or implementation-specific weighting schemes. This paper develops an operator-agnostic axiomatic framework for adaptive weighting in multi-teacher knowledge distillation across three complementary scales: token, task, and context. We formalize structural conditions under which adaptive weighting operators are well-defined, admit multiple non-equivalent implementations, and can be hierarchically composed via product-structure normalization. Within this framework, we establish existence and non-uniqueness of conforming operators, characterize convergence of gradient-based optimization under standard assumptions, analyze stability and perturbation robustness, and provide an abstract formulation of safety-constrained distillation. The results decouple theoretical guarantees from specific weighting formulas, enabling principled analysis of adaptive distillation methods under heterogeneity, distribution shift, and safety constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。