arXiv:2511.02589cs.AI2025-11

测试大模型在真实场景下的计算准确率,发现普遍出错且类型各异。

The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models

  • 用真实生活中的500个任务评估模型计算能力,涵盖金融、物理等领域。
  • 顶尖模型准确率仅45%–63%,主要错误为四舍五入和算术错误。
  • 揭示模型间互补性,适合关注真实计算能力的开发者参考。

我们提出ORCA(AI中量化推理的综合研究)基准,通过Omni计算器引擎验证的输出,评估大语言模型在跨领域真实场景下的定量推理能力。在金融、物理、健康与统计等领域的500个自然语言任务中,五个前沿系统(ChatGPT-5、Gemini 2.5 Flash、Claude Sonnet 4.5、Grok 4、DeepSeek V3.2)准确率仅为45%–63%,其中35%错误源于四舍五入,33%为计算失误。各领域表现显示数学与工程较强,但物理与自然科学较弱。相关性分析(r ≈ 0.40–0.65)表明模型常共同失败,但错误类型不同,体现部分互补性而非冗余。不同于传统数学数据集,ORCA聚焦逐步推理、数值精度与跨域泛化能力。

原文摘要 · Abstract (English)

We present ORCA (Omni Research on Calculation in AI) Benchmark - a novel benchmark that evaluates large language models (LLMs) on multi-domain, real-life quantitative reasoning using verified outputs from Omni's calculator engine. In 500 natural-language tasks across domains such as finance, physics, health, and statistics, the five state-of-the-art systems (ChatGPT-5, Gemini~2.5~Flash, Claude~Sonnet~4.5, Grok~4, and DeepSeek~V3.2) achieved only $45\text{--}63\,\%$ accuracy, with errors mainly related to rounding ($35\,\%$) and calculation mistakes ($33\,\%$). Results in specific domains indicate strengths in mathematics and engineering, but weaknesses in physics and natural sciences. Correlation analysis ($r \approx 0.40\text{--}0.65$) shows that the models often fail together but differ in the types of errors they make, highlighting their partial complementarity rather than redundancy. Unlike standard math datasets, ORCA evaluates step-by-step reasoning, numerical precision, and domain generalization across real problems from finance, physics, health, and statistics.

大模型评估计算准确率真实任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。