指令微调模型在简单词汇限制下会严重失能,响应内容损失高达48%。
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness

- 通过禁止单个标点或常用词,触发模型生成崩溃
- 7个模型响应完整性下降14%至48%,信息量损失超表面表现
- 揭示指令微调导致对表面形式的过度依赖,适合评估与安全研究者
指令微调的大语言模型虽能生成结构化、有帮助的回复,但其鲁棒性如何?我们发现,仅禁止一个标点符号或常见词语,即导致七种不同规模(7B–70B)的指令微调模型响应崩溃,内容完整性损失达14%至48%。10名理工科背景的盲评者确认真实内容缺失,信息类指标退化程度是表面指标的1.5至2.3倍,该结论经超过4,100次自动化成对比较验证(基线偏好77%–100%)。诊断分析表明这是规划失败:双阶段生成可恢复59%–96%的响应长度;提示表示的线性探测器在生成前即可预测响应长度(R²=0.51–0.94),而基础模型则为负值,说明指令微调引入了引发崩溃的表征结构。基础模型在相同约束下无系统性退化,证明指令微调将任务能力绑定于狭窄的表层模板。该效应延伸至实际部署约束(如开场白屏蔽、企业语调规范、法律合规避险、无障碍要求),同样造成22%至34%的性能下降,仅禁用“Certainly!”这一开场词即导致最脆弱模型响应减少40%。此外,标准独立评测仅检测到3.5%质量下降,而成对评测揭示23%下降,暴露当前评估方法的显著盲区。
原文摘要 · Abstract (English)
Instruction-tuned large language models produce helpful, structured responses, but how robust is this helpfulness under trivial constraints? We show that simple lexical constraints (banning a single punctuation character or common word) cause instruction-tuned LLMs to collapse their responses, losing 14--48\% of comprehensiveness across seven models spanning five families (7B--70B, open- and closed-weight). A blinded human evaluation with 10 STEM-trained evaluators confirms genuine content loss, with information criteria degrading $1.5$--$2.3\times$ more than surface criteria, a finding corroborated by over 4,100 automated pairwise comparisons (77--100\% baseline preference) across three LLM judges from two model families. Diagnostic analysis identifies this as a \emph{planning failure}: two-pass generation recovers 59--96\% of response length, and linear probes on prompt representations predict response length with $R^2 = 0.51$--$0.94$ before generation begins. The same probes yield negative $R^2$ on base models, confirming that instruction tuning introduces the representational structure underlying the collapse. Base models show no systematic degradation under identical constraints, demonstrating that instruction tuning couples task competence to narrow surface-form templates. The effect extends to realistic deployment constraints (preamble suppression, corporate tone guidelines, legal compliance hedging, accessibility requirements) causing comparable degradation ($-$22\% to $-$34\%), with suppressing the conversational opener alone (``Certainly!'') causing 40\% collapse on our most fragile model despite restricting only the opening tokens. We further show that standard independent LLM-as-judge evaluation detects only a 3.5\% quality drop where pairwise evaluation reveals 23\%, exposing a methodological blind spot in current evaluation practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。