arXiv:2607.03598cs.CLcs.AI2026-07被引 1

语言模型能准确理解用户意图,但常不按意图回应。

They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It

论文配图:They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
图 1 · 摘自论文原文
  • 用线性探测从隐藏状态解码用户沟通意图,跨模型通用。
  • 6个模型中3个默认不执行意图,但可被定向修复。
  • 修复只需微调特定方向,无需提示词,效果如直接指令。

当用户向语言模型传递信息时,模型常只回应表面内容,而非发送者的真正意图:分享完成的项目,却批评代码;分享深夜潦草的文字,却启动健康检查。本文将发送者的交际意图(格赖斯意义上的‘言外之意’)作为可解释性对象,发现模型具备稳健的意图表征能力,但读出环节存在失败。线性探测可在六种模型、四个系列的基线检查点中,无偏地从默认隐藏状态解码出用户是希望被认可还是被评价。该表征具有泛化性,可延伸至仅需语用推断的意图,以及支持/帮助等语义清晰的意图。行为层面的因果验证基于‘认可/评价’对比,关键在于输出是否响应意图。意图在模型内部可解码的层数早于其驱动输出的层数;不同模型对意图的响应能力存在差异,三类模型表现出这一失败现象,但并非缩放规律。当差距存在时,一个与表征紧密相关的判别方向(在搜索层中找到)可作为因果控制开关:调节它即可恢复正确行为,效果相当于显式指令且无需任何提示。该方向近乎正交于反馈提供轴,因此不是通用反馈旋钮,而是精准路由意图;但在强调节剂量下,其路由的意图甚至可覆盖明确请求。所有结论均通过对照实验验证,显著结果与零结果同等报告。

原文摘要 · Abstract (English)

When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check. We treat the sender's communicative intent, the Gricean what-was-meant, as a first-class interpretability object, and show the failure is one of readout on top of a robust representation. A linear probe decodes the sender's intent, whether they want a thing recognized or evaluated, from a model's default-pass hidden states, cleanly and surface-independently, across six models and four families and in the base checkpoints. The representation generalizes further, to intent that is only pragmatically inferred, and to a second, lexically clean intent (support versus help). The behavioral half of the story, and every causal test, is established on the recognize/evaluate contrast, where what varies is whether the default output acts on the intent. The readout lags the representation in depth within a model (the intent is decodable several layers before it drives the output); across models, which ones act on it by default is model-specific, an observed stratification (three of six show the failure) that we do not read as a scaling law. Where the gap is open, a direction closely tied to the representation, the discriminative direction at a searched-for layer, is a causal handle: steering it recovers the intended behavior, as well as an explicit instruction does and with no prompt at all. This direction is near-orthogonal to the feedback-offering axis, so it routes a represented intent rather than a generic feedback knob, though at the recovery dose the routed intent can override an explicit request. We support each link with controls against obvious deflations and report the nulls as plainly as the confirmations.

意图理解可解释性语言模型因果干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。