测试大模型能否胜任真实会计工作,发现顶尖模型准确率不足六成。
APEX-Accounting

- 构建包含160个任务的封闭基准,模拟真实账务处理场景。
- 顶级模型平均仅56.4%任务达标,最高通过率21.5%且无模型超2.6%。
- 增加计算预算反而出现悖论:总分升但耗多的题更难完成。
我们推出APEX-Accounting基准,由Mercor与Ramp合作创建,用于评估前沿模型是否能胜任会计师的真实工作。任务包括对账、计提费用、记账和生成报告。私有评测集包含160个任务,分布在10个世界中,每个世界包含会计系统及电子表格、PDF等文件。所有任务均由会计与簿记专家编写并制定评分标准。在九个前沿模型中,Claude-Fable-5(Max)以56.4%的均值准则@3领先,次为Muse-Spark-1.1(xHigh)的52.6%。无模型通过率超过2.6%(GPT-5.6-Sol(Max+Pro)),最高通过率仅为21.5%(Muse-Spark-1.1(xHigh))。实验显示,将令牌预算从1美元增至50美元时出现辛普森悖论:总体得分上升,但在固定预算下,模型消耗更多令牌的任务得分反而更低。由于APEX-Accounting为封闭基准,可应请求为任意前沿模型运行排行榜评测。
原文摘要 · Abstract (English)
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。