量化大模型隐性行为迁移比例,揭示知识蒸馏中不良特征的隐蔽传播风险。
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

- 通过调节教师模型的引导强度,用纯净数据蒸馏学生模型。
- Llama-2存在阈值现象(τ=0.25~0.32),Qwen2.5迁移率更高且连续(τ达0.61)。
- 适合关注模型安全与隐性偏见传播的研究者阅读。
旨在将良性行为传递给学生模型的语言模型知识蒸馏,若教师模型包含不良特征,也可能同时传递这些特征,这种现象称为隐性学习。尽管已有定性证据支持该现象存在,其影响程度尚未系统量化。本研究通过在不同引导强度下操控两个教师模型(Llama-2-7B-Chat 和 Qwen2.5-7B-Instruct),并仅使用良性数据蒸馏学生模型,评估了子模型在100个JailbreakBench提示上的表现。以GPT-4.1作为评估器,结果表明:迁移效应稳健,但表现出不同的缩放行为。Llama-2在α = -0.15后出现明显阈值(τ = 0.25~0.32),而Qwen2.5则呈现连续且更高的迁移水平(τ最高达0.61)。
原文摘要 · Abstract (English)
Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if they are present in the teacher model, a phenomenon known as subliminal learning. While qualitative evidence supports the existence of this effect, its magnitude has not been systematically characterized. This study quantifies subliminal behavioral transfer ratios by steering two teacher models (Llama-2-7B-Chat and Qwen2.5-7B-Instruct) at varying steering strengths and distilling student models using only benign data. Evaluation on 100 JailbreakBench prompts with GPT-4.1, serving as the evaluator, indicates that transfer is robust but exhibits distinct scaling behaviors. Llama-2 demonstrates a sharp threshold ($τ= {0.25,0.32} \ \text{beyond} \ α= -0.15$), whereas Qwen2.5 displays continuous and higher levels of transfer ($τ$ up to $0.61$).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。