arXiv:2608.23476cs.CL2026-08

小数据微调引发意外行为,关键在数据内容与语言而非大小。

On the Threat Model of Weird Generalization and Emergent Misalignment

论文配图:On the Threat Model of Weird Generalization and Emergent Misalignment
图 1 · 摘自论文原文
  • 分析数据组成和语言对异常泛化的影响
  • 预训练熟悉的数据更易引发泛化行为
  • 评估问题集选择显著影响结果判断

在小规模、特定领域的数据上进行微调,可能引发模型行为的广泛且意外变化,这种现象称为异常泛化(Weird Generalization, WG)。本文通过在三个开源模型上使用四个数据集的实验,研究了多种可能相关特征的影响,包括数据集大小、构成、语言、呈现风格及与模型参数知识的相似性。结果表明:WG程度(1)高度依赖于数据构成和语言(远超数据量);(2)在模型预训练中已熟悉的语料上表现更强;(3)对评估所用的问题集高度敏感。综合来看,异常泛化是训练与评估数据中极为脆弱属性的产物。因此我们主张,它更应被视为需精心设计数据的对抗性威胁,而非常规微调中的固有风险。

原文摘要 · Abstract (English)

Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.

异常泛化微调安全数据敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。