自动检测语言模型干预后的意外行为变化
Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models

- 对比基线与干预模型在相同提示下的生成结果,识别差异
- 能可靠发现已知行为改变,真实场景中揭示预期与意外变化
- 输出可读的自然语言假设,适合模型审计与安全评估
我们提出一种自动化、对比性的评估流程,用于审计大语言模型干预带来的行为影响。给定基线模型 $M_1$ 与干预模型 $M_2$,该方法在对齐提示上下文中比较两者自由形式、多标记的生成结果,生成人类可读且经过统计验证的自然语言假设,描述模型间的差异,并提炼出重复出现的主题以总结模式。我们在合成设置中注入已知行为变化,证明该流程能可靠恢复这些变化。随后应用于三种真实干预:推理蒸馏、知识编辑与遗忘学习,结果显示该方法能揭示预期及意外的行为转变,区分显著与微弱干预,且在无效应或提示不匹配时不会产生虚假差异。整体上,该流程为干预后模型行为变化提供了统计稳健且可解释的审计工具。
原文摘要 · Abstract (English)
We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model $M_1$ and an intervention model $M_2$, our method compares their free-form, multi-token generations across aligned prompt contexts and produces human-readable, statistically validated natural-language hypotheses describing how the models differ, along with recurring themes that summarize patterns across validated hypotheses. We evaluate the approach in synthetic setting by injecting known behavioral changes and showing that the pipeline reliably recovers them. We then apply it to three real-world interventions, reasoning distillation, knowledge editing and unlearning, demonstrating that the method surfaces both intended and unexpected behavioral shifts, distinguishes large from subtle interventions, and does not hallucinate differences when effects are absent or misaligned with the prompt bank. Overall, the pipeline provides a statistically grounded and interpretable tool for post-hoc auditing of intervention-induced changes in model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。