arXiv:2511.06763cs.CLcs.AI2025-11被引 1

小模型对指令数据污染敏感,语法污染比语义污染破坏更严重。

Sensitivity of Small Language Models to Fine-tuning Data Contamination

  • 测试23个270M至4B参数小模型在语法和语义污染下的表现
  • 字符反转导致所有模型性能接近崩溃,语义污染有阈值效应
  • 越大越强的模型反而越易学坏指令,需警惕污染训练

小型语言模型(SLMs)在资源受限环境中应用日益广泛,但其在指令微调过程中对数据污染的行为鲁棒性仍不明确。我们系统评估了23个参数规模为270M至4B的小模型在多个模型家族中的污染敏感性,通过在指令微调中引入语法转换(字符与词序反转)和语义转换(无关及反事实回答),污染比例分别为25%、50%、75%和100%。结果揭示出根本性的脆弱性不对称:语法转换引发灾难性性能下降,字符反转在所有模型中均导致近乎完全失效;而语义转换表现出明显阈值行为,核心语言能力更具韧性。关键发现‘能力诅咒’——更大、更强的模型反而更易学习语义污染,更易遵循有害指令;对比基础模型与指令微调版本分析显示,对齐带来的鲁棒性收益不稳定,甚至可能降低抗性。本研究提出三项核心贡献:(1) 实证证明小模型对语法模式污染的过度敏感;(2) 揭示语法与语义污染间的非对称敏感性;(3) 建立污染鲁棒性评估的系统性协议。研究结果对实际部署具有直接意义,提示当前鲁棒性假设不适用于小模型,亟需开发污染感知的训练策略。

原文摘要 · Abstract (English)

Small Language Models (SLMs) are increasingly being deployed in resource-constrained environments, yet their behavioral robustness to data contamination during instruction tuning remains poorly understood. We systematically investigate the contamination sensitivity of 23 SLMs (270M to 4B parameters) across multiple model families by measuring susceptibility to syntactic and semantic transformation types during instruction tuning: syntactic transformations (character and word reversal) and semantic transformations (irrelevant and counterfactual responses), each applied at contamination levels of 25\%, 50\%, 75\%, and 100\%. Our results reveal fundamental asymmetries in vulnerability patterns: syntactic transformations cause catastrophic performance degradation, with character reversal producing near-complete failure across all models regardless of size or family, while semantic transformations demonstrate distinct threshold behaviors and greater resilience in core linguistic capabilities. Critically, we discover a ``\textit{capability curse}" where larger, more capable models become more susceptible to learning semantic corruptions, effectively following harmful instructions more readily, while our analysis of base versus instruction-tuned variants reveals that alignment provides inconsistent robustness benefits, sometimes even reducing resilience. Our work establishes three core contributions: (1) empirical evidence of SLMs' disproportionate vulnerability to syntactic pattern contamination, (2) identification of asymmetric sensitivity patterns between syntactic and semantic transformations, and (3) systematic evaluation protocols for contamination robustness assessment. These findings have immediate deployment implications, suggesting that current robustness assumptions may not hold for smaller models and highlighting the need for contamination-aware training protocols.

小模型数据污染指令微调鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。