arXiv:2608.13484cs.CLcs.AI2026-08

大模型知道何时该说‘不知道’,但常强行编造细节。

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

论文配图:Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity
图 1 · 摘自论文原文
  • 用格赖斯合作原则分析大模型的知识边界意识
  • 模型能识别未知实体并预判应答具体程度
  • 适合研究大模型可靠性与对齐机制的学者

当被问及超出知识范围的实体时,大模型通常会编造看似合理的细节,而非采取更安全的泛化表述。我们从格赖斯合作原则出发:不确定指称对象时,合作说话者会退至更泛化的表述层级,以换取真实性。我们探究大模型是否具备实现这一“退让”的能力。基于T-REx的基准测试,考察模型在不同实体熟悉度和指称具体性条件下的表现,提出两个问题:(i) 模型激活值能否编码指称对象是否在知识边界内;(ii) 模型能否预判即将生成的指称具体程度。结果表明两者均成立,但生成阶段未实现协调。模型仍偏好具体回答,即使实体未知且存在正确泛化选项。格赖斯式退让的底层机制存在,但驱动行为的策略缺失。本研究为实现格赖斯对齐迈出第一步,即训练或引导目标使知识边界意识与生成具体性相耦合。

原文摘要 · Abstract (English)

When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark that varies entity familiarity and referent specificity, we probe models to answer two questions: (i) do their activations encode whether a referent falls inside the knowledge boundary, and (ii) do they anticipate the specificity of the referent they are about to generate? We find that the answer to both is yes, but the two signals are not reconciled in generation. Models overwhelmingly prefer specific referents even when the entity is unknown to them, and do so even when offered correct generic alternatives. The substrate for a Gricean retreat is present, but the policy that would act on it is not. We position our findings as a first step toward Gricean alignment, training or steering objectives that couple knowledge-boundary awareness to referent-specificity during generation.

大模型推理知识边界语言模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。