大模型能说会道却未必真懂,成功解释未必能落地。
Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure
- 提出双向一致性悖论:能力与理解脱钩且互逆。
- 低可观测域中成功但误判机制,高可观测域中解释对却无效。
- 需三元框架评估:连贯性、可接地性、解释与行动的关联性。
当智能体能说明某事为何有效时,我们通常视其为真正理解的证据。这假设了有效行为与正确解释相互对应,且一致的解释可靠地反映两者。然而,我指出该假设在当代大语言模型(LLMs)中不成立。我提出‘双向一致性悖论’:在不同认知条件下,能力与可接地性不仅分离,甚至反转。在低可观测领域,LLMs常成功执行任务,却错误识别其成功背后的机制;在高可观测领域,它们常生成准确跟踪可观测因果结构的解释,却无法将诊断转化为有效干预。两种情形下,解释的一致性均保持完整,掩盖了深层的分离。基于编译器优化与超参数调优实验,我构建了‘认识三角形’模型,揭示先验、信号与领域知识在不同可观测性下的交互机制。结果表明,仅凭行为成功或解释准确性,不足以判定理解。我主张评估人工认知主体需采用三元框架:连贯性、可接地性及解释与行动间的恰当奠基关系。大模型中‘知道什么’与‘知道如何’的系统性分离,挑战了来自认识论与当前AI评估实践的既有假设。
原文摘要 · Abstract (English)
When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effective action and correct explanation covary, and that coherent explanation reliably signals both. I argue that this assumption fails for contemporary Large Language Models (LLMs). I introduce what I call the Bidirectional Coherence Paradox: competence and grounding not only dissociate but invert across epistemic conditions. In low-observability domains, LLMs often act successfully while misidentifying the mechanisms that produce their success. In high-observability domains, they frequently generate explanations that accurately track observable causal structure yet fail to translate those diagnoses into effective intervention. In both cases, explanatory coherence remains intact, obscuring the underlying dissociation. Drawing on experiments in compiler optimization and hyperparameter tuning, I develop the Epistemic Triangle, a model of how priors, signals, and domain knowledge interact under varying observability. The results suggest that neither behavioral success nor explanatory accuracy alone suffices for attributing understanding. I argue that evaluating artificial epistemic agents requires a tripartite framework -- coherence, grounding, and a proper basing relation linking explanation to action. The systematic separation of knowing-that and knowing-how in LLMs thus challenges assumptions inherited from both epistemology and current AI evaluation practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。