提出因果中介方法,消除对话主题对神经网络解释的干扰。
Removing Spurious Correlation from Neural Network Interpretations
- 引入因果中介分析,控制话题等混杂因素影响
- 调整话题影响后,毒性判断的局部化程度降低
- 适用于需可信解释的大模型可解释性研究
现有识别神经元中不良行为的方法未考虑话题等混杂因素的影响。本文表明,混杂因素会导致虚假相关,并提出一种新的因果中介方法以控制话题的影响。在两个大语言模型上的实验验证了定位假说,结果显示,在调整对话话题效应后,毒性特征的局部化程度下降。
原文摘要 · Abstract (English)
The existing algorithms for identification of neurons responsible for undesired and harmful behaviors do not consider the effects of confounders such as topic of the conversation. In this work, we show that confounders can create spurious correlations and propose a new causal mediation approach that controls the impact of the topic. In experiments with two large language models, we study the localization hypothesis and show that adjusting for the effect of conversation topic, toxicity becomes less localized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。