模型能识别逻辑对错,但未必真懂,输出常不反映内部理解。
When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

- 用匹配的真假前提-结论对测试五种开源模型的逻辑验证能力
- 逻辑正确性在隐藏状态中几乎完全可解码,即使行为错误也成立
- 内部有逻辑表示却难影响输出,说明理解与表达是两回事
大型语言模型看似具备逻辑推理能力,但仅凭答案对错无法揭示其内部表征。我们通过匹配的有效-无效前提-结论对,在五种开源Transformer模型中研究逻辑验证能力,涵盖不同推理类型、语义领域、模板结构和难度水平。尽管行为表现接近随机,逻辑有效性仍几乎完全可从隐藏状态中解码,并在未见模板、领域和推理家族下保持强可解码性。即使在行为错误样本中,有效性依然高度可解码(在正确性条件定义明确的情况下)。然而,全面的留一法测试揭示了泛化局限,沿探测所得有效性方向的干预效果微弱且不具特异性,远低于随机对照组。结果表明:表示有效性、表达有效性、因果使用有效性三者独立。逻辑相关信息可在隐藏状态中强解码,却未必可靠地体现在输出中。
原文摘要 · Abstract (English)
Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models using matched valid--invalid premise--claim pairs that vary across inference families, semantic domains, templates, and difficulty levels. Despite near-chance behavioral performance, logical validity is often almost perfectly decodable from hidden states and remains strongly decodable under held-out templates, domains, and inference families. Validity also remains highly decodable on behaviorally incorrect examples in the conditions where correctness-conditioned evaluation is well defined. At the same time, exhaustive leave-one-out tests reveal clear limits to this generalization, and interventions along probe-derived validity directions have only weak, nonspecific effects compared with random controls. Our results suggest that representing validity, expressing it in behavior, and using it causally are distinct. Validity related information can be strongly decodable from a model's hidden states without being reliably expressed in its output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。