评测大模型代理在金融电子表格全流程任务中的表现
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

- 设计三维度评估体系:准确度、公式逻辑、格式规范
- 18个代理中最强者仍不达专业金融标准,复杂任务性能骤降
- 适合关注AI办公自动化与金融建模可信度的研究者
大型语言模型代理正被期待完成从高阶指令到完整成果的端到端工作流。为满足企业需求,前沿人工智能实验室已开发出能从零构建完整电子表格的代理,尤其在金融领域,财务建模、预测和情景分析等核心流程均依赖电子表格。然而现有基准仅聚焦问答或单公式修改,无法衡量此新能力。为此,我们提供了首个针对端到端电子表格任务的评估,重点关注经济关键型金融工作流。由于交付成果常需多利益相关方审核与修订,质量判断需涵盖可读性、易修改性等高层标准。我们提出包含准确性、公式逻辑、格式规范三个维度的评估体系,每项细化为符合行业标准的子指标。对18个代理的评估显示,即使最强代理也未达到基本专业金融标准,且在超过少量链式计算后性能显著下降。这表明当前代理尚无法在现实工作流所需复杂度下可靠生成专业级电子表格。
原文摘要 · Abstract (English)
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs have developed agents that can construct entire spreadsheets from scratch. This is especially relevant in finance, where core workflows such as financial modeling, forecasting, and scenario analysis are commonly conducted through spreadsheets. Yet, existing spreadsheet benchmarks do not measure this new capability, focusing instead on question-answering or single-formula edits. To address this gap, we provide one of the first evaluations of agents on end-to-end spreadsheet tasks, focusing on economically critical financial workflows such as modeling and scenario analysis. Since deliverables therein are routinely reviewed and revised by multiple stakeholders, judging their quality necessarily involves high-level criteria such as readability or ease of modification. To reflect the multidimensional nature of solution quality, we develop an evaluation taxonomy comprising three dimensions: Accuracy, Formula, and Format, each comprising fine-grained criteria that reflect professional standards. Evaluating over 18 agents, the benchmark reveals that even the strongest agents fall short of basic professional finance standards, and their performance degrade sharply as the difficulty increases beyond a few chained calculations. This suggests that current agents are not yet able to reliably produce professional-quality spreadsheets at the level of complexity real-world workflows demand.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。