arXiv:2510.06700cs.CL2025-10ACL被引 4

语言模型会把内容合理性误当成逻辑正确性,研究揭示了其内部表征机制。

How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects

  • 通过分析模型内部表征,发现逻辑有效性和内容合理性高度对齐。
  • 用定向向量操控后,可显著改变判断结果,证明二者可相互影响。
  • 提出去偏方法,能降低混淆,提升模型推理准确性,适合研究逻辑偏差者。

人类和大型语言模型(LLMs)均存在内容效应:推理问题语义内容的合理性会影响对其逻辑有效性的判断。尽管人类的这种现象可用双过程推理理论解释,但语言模型中的机制仍不明确。本文通过研究语言模型如何在其内部表征中编码有效性与合理性,发现这两个概念在表征几何中呈线性且高度对齐,导致模型将合理性误判为有效性。利用控制向量,我们证明合理性向量可因果性地影响有效性判断,反之亦然;且二者对齐程度可预测不同模型的行为内容效应大小。最后,我们构建去偏向量,实现概念解耦,有效降低内容效应并提升推理准确率。研究深化了对抽象逻辑概念在语言模型中表征的理解,并指出表征干预是实现更严谨逻辑系统的重要路径。

原文摘要 · Abstract (English)

Both humans and large language models (LLMs) exhibit content effects: biases in which the plausibility of the semantic content of a reasoning problem influences judgments regarding its logical validity. While this phenomenon in humans is best explained by the dual-process theory of reasoning, the mechanisms behind content effects in LLMs remain unclear. In this work, we address this issue by investigating how LLMs encode the concepts of validity and plausibility within their internal representations. We show that both concepts are linearly represented and strongly aligned in representational geometry, leading models to conflate plausibility with validity. Using steering vectors, we demonstrate that plausibility vectors can causally bias validity judgements, and vice versa, and that the degree of alignment between these two concepts predicts the magnitude of behavioral content effects across models. Finally, we construct debiasing vectors that disentangle these concepts, reducing content effects and improving reasoning accuracy. Our findings advance understanding of how abstract logical concepts are represented in LLMs and highlight representational interventions as a path toward more logical systems.

语言模型逻辑推理表征分析去偏方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。