SeqKD让学生在未见原始数据时仍会记忆大量内容,需警惕幻觉风险。
Memorization Inheritance in Sequence-Level Knowledge Distillation for Neural Machine Translation
- 学生通过序列级知识蒸馏间接继承教师的记忆模式。
- 学生对原数据的精确匹配率3.4%,可提取记忆率达57%。
- 新方法Adaptive-SeqKD可有效降低记忆与幻觉现象,适合关注模型安全的开发者。
本文研究序列级知识蒸馏(SeqKD)中,学生模型如何继承教师模型在实例级别上的记忆行为。尽管学生未直接接触原始训练数据,其记忆程度仍高于同规模基准模型——精确匹配率为3.4%,可提取记忆率达57%,且幻觉率显著上升。进一步分析发现,在低质量数据子集和特定反事实记忆(CM)得分子集上,学生表现出更强的去噪能力。为此,本文提出改进方案Adaptive-SeqKD,通过干预蒸馏过程减少记忆与幻觉。总体而言,使用SeqKD需保持警惕:学生不仅继承教师性能优势,也继承其错误模式,必须进行主动监控。
原文摘要 · Abstract (English)
In this work, we explore how instance-level memorization in the teacher Neural Machine Translation (NMT) model gets inherited by the student model in sequence-level knowledge distillation (SeqKD). We find that despite not directly seeing the original training data, students memorize more than baseline models (models of the same size, trained on the original data) -- 3.4% for exact matches and 57% for extractive memorization -- and show increased hallucination rates. Further, under this SeqKD setting, we also characterize how students behave on specific training data subgroups, such as subgroups with low quality and specific counterfactual memorization (CM) scores, and find that students exhibit amplified denoising on low-quality subgroups. Finally, we propose a modification to SeqKD named Adaptive-SeqKD, which intervenes in SeqKD to reduce memorization and hallucinations. Overall, we recommend caution when applying SeqKD: students inherit both their teachers' superior performance and their fault modes, thereby requiring active monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。