首个评估语言模型上下文使用解释的基准框架,可精准定位模型依赖的上下文。
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
- 构建可控测试用例,以真实上下文使用情况为黄金标准评估解释效果。
- 机械解释方法MechLight在四种场景下表现最优,但长上下文仍难准确解释。
- 适合关注大模型可解释性、需验证上下文依赖关系的研究者使用。
语言模型在生成回应时是否利用了提供的上下文信息,对用户而言仍不透明,无法判断模型是依赖参数记忆还是上下文内容,也无法识别具体影响输出的上下文片段。亮点解释(HEs)通过标记具体上下文片段和词元,可自然解决此问题。然而,现有工作缺乏对HE有效性的直接评估。本文提出首个基于已知真实上下文使用情况的黄金标准评估框架,克服了间接代理评估的局限性。为验证框架普适性,我们评估了四种HE方法——三种已有技术及我们适配的机械可解释性方法MechLight——在四个上下文场景、四个数据集和五种语言模型上的表现。结果表明,MechLight在所有场景中表现最佳;但所有方法在处理长上下文时均表现不佳,并存在位置偏差,揭示出解释准确性面临根本性挑战,亟需新方法实现大规模可靠上下文使用解释。
原文摘要 · Abstract (English)
Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users, who cannot determine whether models draw from parametric memory or provided context, nor identify which specific context pieces inform the response. Highlight explanations (HEs) offer a natural solution as they can point the exact context pieces and tokens that influenced model outputs. However, no existing work evaluates their effectiveness in accurately explaining context utilisation. We address this gap by introducing the first gold standard HE evaluation framework for context attribution, using controlled test cases with known ground-truth context usage, which avoids the limitations of existing indirect proxy evaluations. To demonstrate the framework's broad applicability, we evaluate four HE methods -- three established techniques and MechLight, a mechanistic interpretability approach we adapt for this task -- across four context scenarios, four datasets, and five LMs. Overall, we find that MechLight performs best across all context scenarios. However, all methods struggle with longer contexts and exhibit positional biases, pointing to fundamental challenges in explanation accuracy that require new approaches to deliver reliable context utilisation explanations at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。