大模型看不懂数学算法执行,预测能力远不如随机猜测。
Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction
- 用算法执行结果做因果发现测试,检验模型推理能力
- 八款主流大模型预测区间普遍过宽且多数不包含真实均值
- 最佳模型仅小幅提升,疑似记忆数据而非真正推理
大型语言模型(LLMs)展现出惊人的知识广度,但其对计算过程的推理能力仍不明确。这一差距对依赖模型进行算法选择与部署的实践者至关重要。本文以因果发现为测试场景,评估八款前沿大模型在基于算法执行所得真实值上的表现。结果发现,各模型普遍存在系统性、近乎完全的失败:预测范围远超真实置信区间,却仍无法覆盖真实算法均值。多数模型表现甚至劣于随机猜测。最优模型的微弱改进暗示其可能依赖基准记忆而非系统性推理。我们称此现象为‘算法盲视’,表明模型在算法的陈述性知识与校准后的程序性预测之间存在根本性鸿沟。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate remarkable breadth of knowledge, yet their ability to reason about computational processes remains poorly understood. Closing this gap matters for practitioners who rely on LLMs to guide algorithm selection and deployment. We address this limitation using causal discovery as a testbed and evaluate eight frontier LLMs against ground truth derived from algorithm executions. We find systematic, near-total failure across models. The predicted ranges are far wider than true confidence intervals yet still fail to contain the true algorithmic mean in most cases. Most models perform worse than random guessing. The best model's marginal improvement points to benchmark memorization rather than principled reasoning. We term this failure algorithmic blindness and argue it reflects a fundamental gap between declarative knowledge about algorithms and calibrated procedural prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。