arXiv:2602.01378cs.CLcs.AI2026-02被引 1

提出新方法RIS,让大模型解释更稳定可靠。

Context Dependence and Reliability in Autoregressive Language Models

  • 用新评分法分离上下文关键信息与冗余内容
  • 实验显示其解释稳定性显著优于传统方法
  • 适合关注模型可解释性与安全性的研究者

大型语言模型(LLMs)在生成输出时依赖大量上下文,其中常包含提示、检索内容和交互历史等冗余信息。在关键应用中,识别真正影响输出的上下文元素至关重要,而现有解释方法因冗余和重叠信息难以准确判断。输入微小变化可能导致归因分数剧烈波动,削弱可解释性并引发如提示注入等风险。本文提出RIS(Redundancy-Insensitive Scoring of Explanation),量化每个输入相对于其他输入的唯一影响,降低冗余干扰,提供更清晰、稳定的归因结果。实验表明,RIS比传统方法具有更强的鲁棒性,凸显条件信息对可信解释与监控的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) generate outputs by utilizing extensive context, which often includes redundant information from prompts, retrieved passages, and interaction history. In critical applications, it is vital to identify which context elements actually influence the output, as standard explanation methods struggle with redundancy and overlapping context. Minor changes in input can lead to unpredictable shifts in attribution scores, undermining interpretability and raising concerns about risks like prompt injection. This work addresses the challenge of distinguishing essential context elements from correlated ones. We introduce RISE (Redundancy-Insensitive Scoring of Explanation), a method that quantifies the unique influence of each input relative to others, minimizing the impact of redundancies and providing clearer, stable attributions. Experiments demonstrate that RISE offers more robust explanations than traditional methods, emphasizing the importance of conditional information for trustworthy LLM explanations and monitoring.

可解释性大模型归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。