评测大模型处理真实企业电子表格全流程的能力,发现当前系统准确率不足三成。
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

- 构建基于真实业务数据的全流程电子表格评测基准
- 多表跨页依赖任务平均需593次修改,最高准确率仅34.89%
- 适合研究自动化办公与大模型应用的开发者参考
电子表格广泛用于商业分析、财务建模、报告生成和决策制定。然而,现有评测多聚焦单一公式生成或局部单元格编辑,无法反映真实业务场景中的端到端工作流。我们提出 extsc{SpreadsheetBench 2},一个面向电子表格智能体的工作流级评测基准,涵盖生成、调试和可视化三类任务。基准基于真实企业数据(包括财务报告与公司备案文件),经领域专家标注与验证,共包含321个任务,每项任务平均涉及11.8个工作表,需进行593.5次单元格修改,体现复杂多表结构与跨表依赖关系。我们在统一多轮代理框架下评估八个前沿大语言模型,并引入多个基于LLM的电子表格产品作为补充基线。结果表明,当前系统在真实工作流中仍不可靠:最优模型整体任务准确率为34.89%,调试准确率低至12.00%。轨迹分析与失败归因显示,检查不充分与目标单元格误选是主要瓶颈。这些发现确立了 extsc{SpreadsheetBench 2}作为推进可靠电子表格自动化的挑战性测试平台。
原文摘要 · Abstract (English)
Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations such as single-formula generation or local cell edits, and therefore fail to capture end-to-end workflows in realistic business settings. We introduce \textsc{SpreadsheetBench 2}, a workflow-level benchmark for spreadsheet agents that covers three task categories: generation, debugging, and visualization. The benchmark is constructed from authentic business data, including financial reports and corporate filings, and is annotated and validated by domain experts. The benchmark contains 321 tasks; each instance averages 11.8 worksheets and requires 593.5 cell modifications, reflecting large multi-sheet workbooks with cross-sheet dependencies. We evaluate eight frontier large language models under a unified multi-turn agent scaffold, and additionally include several LLM-based spreadsheet products as complementary baselines. Results show that current systems remain far from reliable on real-world workflows: the best model achieves 34.89\% overall task accuracy, and debugging accuracy is as low as 12.00\%. Trajectory analysis and a failure taxonomy further indicate that insufficient spreadsheet inspection and incorrect target-cell selection are the dominant bottlenecks. Together, these findings position \textsc{SpreadsheetBench 2} as a challenging testbed for advancing reliable spreadsheet automation. Project page: https://spreadsheetbench.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。