arXiv:2608.27756cs.CL2026-08

用删除法测大模型是否真依赖上下文,发现它常在缺关键信息时仍瞎答。

Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning

论文配图:Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning
图 1 · 摘自论文原文
  • 通过删掉谜题中的单个上下文例句,测试模型对信息的依赖程度
  • 三款前沿大模型在关键上下文被删后仍常正确作答,说明它们不靠上下文
  • 可用来研究模型推理机制,适合关注模型可信性的研究人员

判断大语言模型是基于上下文还是先验知识得出答案,仍是根本性挑战。自洽的语言奥数谜题提供了一个受控环境,所有答案仅来自专家设计的上下文示例,无需外部知识。删除单个上下文例句可消除特定问题所需信息,同时保持其他部分不变。我们利用这一特性提出诊断框架,分析单个上下文例句的作用。基于53道英国语言学奥林匹克竞赛题目,生成两种修改版本:(1)均匀随机删除;(2)目标删除(受纠错码启发),移除唯一承载必要信息的结构性关键例句。通过“问题损伤度”量化影响,分类谜题为脆弱或稳健。在指令要求信息不足时拒绝回答的前提下,评估三款前沿大模型,发现它们极少拒绝回答,即使在关键上下文被移除后仍常给出正确答案。这些发现促使进一步探究上下文推理、先验知识、记忆与语言推断。该框架不仅支持拒绝机制研究,还可实现细粒度分析,包括因果干预、停止集分析、目标污染研究及机制可解释性。

原文摘要 · Abstract (English)

Determining whether large language models derive answers from context or prior knowledge remains a fundamental challenge. Self-contained linguistic olympiad puzzles provide a controlled setting where all answers derive solely from expert-designed context examples without external knowledge. Removing individual context examples can eliminate information needed for specific questions while leaving the rest of the puzzle unchanged. We leverage this to introduce a diagnostic framework for analyzing individual context examples. Using 53 UK Linguistics Olympiad puzzles, we generate two modified variants by deleting a single context example: (1) uniform random deletion, and (2) targeted deletion (inspired by error-correcting codes) to remove a structurally load-bearing example uniquely carrying necessary information. We formalize this impact using a Question Damage Score to classify puzzles as fragile or robust. Evaluating three frontier LLMs under instructions to abstain when information is insufficient, we find they rarely abstain, often continuing to produce correct answers after load-bearing context is removed. These findings motivate further investigation into context-based reasoning, prior knowledge, memorization, and linguistic inference. Beyond abstention, the framework enables fine-grained analyses of context reliance, including causal interventions, stopping-set analysis, targeted contamination studies, and mechanistic interpretability.

模型可解释性上下文依赖语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。