人类难以看懂AI推理文本中的因果关系,存在严重误解。
Humans Perceive Wrong Narratives from AI Reasoning Texts
- 用反事实测试评估人类对推理步骤因果影响的判断能力
- 人类准确率仅29%,远低于模型实际因果结构
- 提示:适合关注AI可解释性与人机认知差异的研究者
新一代AI模型在生成答案前会输出逐步推理文本,看似提供了可读的计算过程窗口,被广泛用于透明性和可解释性。然而,人类对这些文本的理解是否真实反映模型的实际计算过程尚不明确。本文通过反事实测量设计问题,检验人类识别推理步骤间因果影响的能力。结果发现,参与者准确率仅为29%,仅略高于随机水平(25%),即使在高一致性问题中多数投票准确率也仅达42%。这揭示了人类对推理文本的理解与模型实际使用方式之间存在根本性脱节,挑战了其作为简单可解释性工具的有效性。我们主张将推理文本视为需深入研究的产物,而非表面可信的解释,并强调理解模型非人类的语言使用方式是关键研究方向。
原文摘要 · Abstract (English)
A new generation of AI models generates step-by-step reasoning text before producing an answer. This text appears to offer a human-readable window into their computation process, and is increasingly relied upon for transparency and interpretability. However, it is unclear whether human understanding of this text matches the model's actual computational process. In this paper, we investigate a necessary condition for correspondence: the ability of humans to identify which steps in a reasoning text causally influence later steps. We evaluated humans on this ability by composing questions based on counterfactual measurements and found a significant discrepancy: participant accuracy was only 29%, barely above chance (25%), and remained low (42%) even when evaluating the majority vote on questions with high agreement. Our results reveal a fundamental gap between how humans interpret reasoning texts and how models use it, challenging its utility as a simple interpretability tool. We argue that reasoning texts should be treated as an artifact to be investigated, not taken at face value, and that understanding the non-human ways these models use language is a critical research direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。