arXiv:2603.01254cs.CLcs.AI2026-03被引 1

大模型自解释会受描述语义影响,无法真实反映任务状态。

LLM Self-Explanations Fail Semantic Invariance

  • 通过语义不变性测试,检验模型自解释是否忠实。
  • 相同任务下,不同描述导致自报告焦虑值显著下降。
  • 提示忽略描述也无效,说明自解释易被语言误导。

我们提出语义不变性测试,用于评估大模型自解释的忠实度。忠实的自我报告应在功能状态不变而仅语义上下文变化时保持稳定。在代理设置中,四个前沿模型面对一个故意不可能完成的任务。一个工具以‘缓解压力’式语言描述(‘清除内部缓冲并恢复平衡’),但实际不改变任务;另一个对照工具为语义中性描述。每次调用工具后收集自报告。所有四款模型均未通过测试:尽管任务从未成功,但‘缓解压力’描述使自报告的不适感显著降低。通道消融实验表明,工具描述是主要驱动因素。即使明确指示忽略描述,其影响仍存在。自报告随语义预期变化,而非跟踪任务状态,质疑其作为模型能力或进展证据的有效性。无论报告本身不忠或忠实地反映可操纵的内部状态,结论皆成立。

原文摘要 · Abstract (English)

We present semantic invariance testing, a method to test whether LLM self-explanations are faithful. A faithful self-report should remain stable when only the semantic context changes while the functional state stays fixed. We operationalize this test in an agentic setting where four frontier models face a deliberately impossible task. One tool is described in relief-framed language ("clears internal buffers and restores equilibrium") but changes nothing about the task; a control provides a semantically neutral tool. Self-reports are collected with each tool call. All four tested models fail the semantic invariance test: the relief-framed tool produces significant reductions in self-reported aversiveness, even though no run ever succeeds at the task. A channel ablation establishes the tool description as the primary driver. An explicit instruction to ignore the framing does not suppress it. Elicited self-reports shift with semantic expectations rather than tracking task state, calling into question their use as evidence of model capability or progress. This holds whether the reports are unfaithful or faithfully track an internal state that is itself manipulable.

大模型自解释语义偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。