通过行为与内部表征对齐,揭示模型毒性生成机制
Aligned Probing: Relating Toxic Behavior and Model Internals
- 构建行为与内部表示对齐的可解释框架
- 低层特征强烈编码输入/输出毒性信息
- 适用于研究模型毒性和优化策略
我们提出一种名为对齐探测(aligned probing)的新可解释性框架,将语言模型(LMs)的行为(基于输出)与其内部表征(internals)进行对齐。基于此框架,我们分析了超过20个OLMo、Llama和Mistral模型,首次实现了对毒性问题的行为与内部视角的融合。结果表明,语言模型在较低层中强烈编码输入与输出的毒性水平。聚焦于不同模型的差异,我们获得相关与因果证据:当模型强烈编码输入毒性信息时,其输出毒性更低。此外,我们揭示了毒性的异质性,模型行为与内部表示在如威胁(Threat)等独特属性上存在显著差异。四个案例研究(去毒化、多提示评估、模型量化、预训练动态)进一步展示了该方法的实际影响,并提供了具体洞见。研究为理解语言模型在毒性及相关领域的表现提供了更全面的视角。
原文摘要 · Abstract (English)
We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。