测试大模型掷骰子概率推理能力,发现其在反直觉问题上表现差。
How reliable are LLMs when it comes to playing dice?

- 构建标准与反直觉概率题集,对比大模型表现。
- 标准题准确率96%,反直觉题仅59%。
- 提示词伪装或误导会降低34%性能,无模型免疫。
我们通过受控基准测试研究大语言模型在离散概率问题上的概率推理能力。构建了两套数据集:一套为标准习题,另一套为旨在触发启发式推理的反直觉习题,并评估了8个主流模型,每种模型均在有无思维链提示下进行测试。模型在标准问题上的平均准确率为0.96,但在反直觉问题上仅为0.59。进一步实证发现存在令牌偏差:当使用非标准表述替代标准形式时,性能下降超过20%。在提示中嵌入误导性建议可使性能下降达34%,且无一模型能免疫。综合来看,当前大语言模型并非真正的概率推理者,尽管它们在复杂数学问题上表现优异。
原文摘要 · Abstract (English)
We investigate the probabilistic reasoning capabilities of large language models through a controlled benchmarking study on discrete probability problems. We constructed two datasets, respectively a set of standard exercises and a set of counterintuitive exercises, designed to trigger heuristic reasoning, and evaluated 8 state-of-the-art models, each tested with and without Chain-of-Thought prompting. Models achieve an average accuracy of 0.96 on standard problems but only 0.59 on counterintuitive ones. We further provide empirical evidence of token bias: performance drops by over 20% when canonical formulations are replaced by disguised variants. Embedding misleading suggestions in the prompt reduces performance by up to 34%, with no model proving immune. Taken together, the reported findings suggest that current LLMs are not yet genuine probabilistic reasoners, despite their success in advanced mathematical problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。