改进遗忘探针中信息移除方式,提升因果干预准确性
Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE
- 用均值投影或LEACE替代INLP进行信息移除
- 新方法减少对非目标信息的干扰,提升干预精准度
- 适合研究模型内部表征与语言信息关系的学者
遗忘探针是一种用于检验特定语言信息对模型行为影响的技术,通过识别并移除相关信息,评估主任务性能是否变化。若移除信息相关,则性能应下降。该方法难点在于仅移除目标信息而保持其他信息不变。已有研究表明,广泛使用的迭代零空间投影(INLP)在消除目标信息时会引入随机表示修改。本文证明,均值投影(MP)和LEACE两种新方法能更精准地移除信息,从而增强通过遗忘探针获得行为解释的潜力。
原文摘要 · Abstract (English)
Amnesic probing is a technique used to examine the influence of specific linguistic information on the behaviour of a model. This involves identifying and removing the relevant information and then assessing whether the model's performance on the main task changes. If the removed information is relevant, the model's performance should decline. The difficulty with this approach lies in removing only the target information while leaving other information unchanged. It has been shown that Iterative Nullspace Projection (INLP), a widely used removal technique, introduces random modifications to representations when eliminating target information. We demonstrate that Mean Projection (MP) and LEACE, two proposed alternatives, remove information in a more targeted manner, thereby enhancing the potential for obtaining behavioural explanations through Amnesic Probing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。