arXiv:2603.14761cs.AIcs.CL2026-03

测试大模型常识推理能力,发现顶尖模型仍有近20%错误。

BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models

  • 设计100道脑筋急转弯题,覆盖20类常识推理陷阱
  • 顶级模型准确率仅80.3%,最差的仅39.7%
  • 跨语言测试显示模型推理缺陷非语言特有

大型语言模型在标准基准上表现优异,却常无法回答人类秒答的问题。我们提出BrainBench,一个包含100道脑筋急转弯题的基准,涵盖20个精心设计的类别,每类针对一类常识推理失效模式,如隐含物理约束、语义范围陷阱和默认假设干扰。评估八款前沿模型——四款Claude系列与四款GPT系列——采用零样本协议,每题进行10次独立运行。最佳模型Claude Opus 4.6(扩展思维)准确率为80.3%,最差模型GPT-4o仅为39.7%。即使顶尖模型,准确率与一致性之间仍存在6-16个百分点差距,揭示其推理具有随机性。中文跨语言测试显示多数模型性能下降2-8个百分点,证实这些失败反映的是推理缺陷而非语言特异性。BrainBench提供细粒度诊断工具,可识别模型何时用表面启发式替代真正常识推理。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve impressive scores on standard benchmarks yet routinely fail questions that any human would answer correctly in seconds. We introduce BrainBench, a benchmark of 100 brainteaser questions spanning 20 carefully designed categories, each targeting a specific commonsense reasoning failure mode in LLMs. Categories range from implicit physical constraints ("Should I walk or drive my rental car to the return lot?") to semantic scope tricks and default assumption hijacks. We evaluate eight frontier models -- four from the Claude family and four from the GPT family -- using a zero-shot protocol with 10 independent runs per question. The best model, Claude Opus 4.6 with extended thinking, achieves only 80.3% accuracy; the worst, GPT-4o, scores 39.7%. Even top-performing models exhibit a 6-16 percentage-point gap between accuracy and consistency, revealing stochastic reasoning. Cross-lingual evaluation in Chinese shows most models degrade by 2-8 percentage points, confirming that these failures reflect reasoning deficits rather than language-specific artifacts. BrainBench provides a fine-grained diagnostic tool for identifying where and why LLMs substitute surface heuristics for genuine commonsense reasoning.

常识推理大模型评测脑筋急转弯认知缺陷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。