arXiv:2604.03192cs.CLcs.AI2026-04被引 1

通过可靠性感知机制提升低资源摘要生成效果

Reliability Gated Multi-Teacher Distillation for Low Resource Abstractive Summarization

  • 根据教师间一致性动态分配监督信号,优化知识蒸馏
  • 在多种语言下实现3.2倍压缩,保留71%-122%的摘要质量
  • 揭示单裁判评估的校准偏差,适合多教师蒸馏研究者

我们从可靠性视角研究低资源抽象摘要中的多教师知识蒸馏。提出EWAD(熵加权一致感知蒸馏)机制,基于教师间一致性在教师蒸馏与真实标签间动态分配监督信号;引入CPDP(容量比例发散保持)约束,控制学生模型相对于异构教师的位置关系。在两个孟加拉语数据集、13个BanglaT5变体及8个Qwen2.5实验中发现,对数层蒸馏带来最可靠的性能提升;更复杂的蒸馏虽提升短摘要语义相似度,但损害长输出质量。跨语言伪标签蒸馏在十种语言上实现3.2倍压缩,保留71-122%的教师ROUGE-L得分。人工验证的多评委大模型评估进一步揭示了单评委流水线中的校准偏差。总体表明,可靠性感知蒸馏有助于判断多教师监督何时有效,以及数据扩展何时优于损失工程。

原文摘要 · Abstract (English)

We study multiteacher knowledge distillation for low resource abstractive summarization from a reliability aware perspective. We introduce EWAD (Entropy Weighted Agreement Aware Distillation), a token level mechanism that routes supervision between teacher distillation and gold supervision based on inter teacher agreement, and CPDP (Capacity Proportional Divergence Preservation), a geometric constraint on the student position relative to heterogeneous teachers. Across two Bangla datasets, 13 BanglaT5 ablations, and eight Qwen2.5 experiments, we find that logit level KD provides the most reliable gains, while more complex distillation improves semantic similarity for short summaries but degrades longer outputs. Cross lingual pseudo label KD across ten languages retains 71-122 percent of teacher ROUGE L at 3.2x compression. A human validated multi judge LLM evaluation further reveals calibration bias in single judge pipelines. Overall, our results show that reliability aware distillation helps characterize when multi teacher supervision improves summarization and when data scaling outweighs loss engineering.

知识蒸馏低资源摘要生成可靠性感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。