发现大模型在上下文与训练数据重叠时,难以区分记忆与当前输入的贡献。
How Context Attribution Handles What the Model Already Knows

- 提出新评估协议和基准数据集,量化上下文归因可靠性
- 实验显示四种方法在知识重叠时产生不可靠得分
- 证明现有方法无法有效分离权重记忆与上下文学习
大语言模型的上下文归因方法旨在识别输入上下文对模型输出的贡献。近期研究虽取得初步成功,但我们发现当上下文与训练数据重叠时,这些方法无法区分上下文贡献与权重内记忆(IW)贡献,导致归因分数不可靠。为此,本文提出:1)基于四项新指标(基础模型上下文归因得分BCS、跨模型归因一致性CAC、归因保留率APS、源分离精确度SSP)的评估协议;2)包含真实来源标签的基准数据集WMDP-Cyber++,以系统评估存在权重重叠时的归因表现。在四种主流归因方法上的实验表明,当上下文知识存在于模型权重中时,这些方法提供的是不忠实的归因结果。进一步地,我们尝试利用归因得分进行源分离(IW vs. ICL),但结果表明现有方法无法实现有效解耦。
原文摘要 · Abstract (English)
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe that when the context overlaps with the training data, these methods can- not disentangle in-context from in-weight (IW) contributions, producing unreliable scores. Based on this observation, in this work, we introduce: 1) an evaluation protocol that relies on four new metrics (base-model context attribution score (BCS), cross-model context attribution consistency (CAC), attribution preservation score (APS), source separation pre- cision (SSP)) and 2) a benchmark dataset (WMDP-Cyber++) with ground-truth provenance labels to systematically assess attribution under IW overlap. In our experiments across four well-known context attribution methods, we demonstrate that they provide unfaithful attribution when the knowledge from the context also exists in the weights. Finally, we adapt these methods for source separation (IW vs. in-context learning (ICL)) and show that they cannot do the disentanglement based on the contributive score
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。