评测大模型在金融表格任务中的表现,发现其准确率不足50%。
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

- 构建金融领域真实场景的复杂表格任务集,含3225条评分标准。
- 顶尖大模型平均得分低于50%,动态正确性是主要短板。
- 提供开源评估框架与人类验证的评分体系,适合模型评测者使用。
我们提出BlueFin,一个针对大型语言模型(LLM)代理在专业金融领域中处理表格文件的合成、操作与理解任务的基准测试。尽管全球付费表格软件用户达数亿,远超专业开发者数量,但针对该领域的模型能力研究仍严重不足。为此,我们整理了131个具有现实意义的复杂任务,包含3,225项细粒度评分标准;这些标准由专家团队验证,确保了高质量、可信赖的评估结果。我们的语言模型裁判在宏平均F1上达到0.839,与专家共识一致性高达α=0.826。前沿大模型在此基准上表现不佳,最强模型在所有任务上的平均得分不足50%,尤其在动态正确性方面存在明显缺陷。本研究贡献包括三个类别共3225项任务的数据集、开源的评估工具链以及对当前先进模型性能的系统性分析。
原文摘要 · Abstract (English)
We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet workbooks in the professional finance domain. Though estimates of the global population of paying users of spreadsheet software range in the hundreds of millions -- an order of magnitude more than the estimated global population of professional developers -- comparatively fewer resources have been devoted to exploring and expanding LLM capabilities in the spreadsheet domain, with fewer still dedicated to mirroring real occupational tasks encountered by those in professional finance roles. In response, we curate a set of 131 challenging, complex tasks with real-world relevance in the domain, containing 3,225 granular rubric criteria; notably, our rubric criteria and LM judge evaluations are validated by a team of expert human annotators, resulting in high-quality, granular evaluations of complex tasks that are difficult to verify programmatically but can be reliably evaluated by an LM judge agent. Our judge achieves parity with expert consensus ($α=0.826$) with a macro-F1 score of 0.839. Frontier LLMs demonstrate poor performance on the challenging benchmark, with the strongest LLMs achieving less than 50\% average scores across tasks -- models exhibit particular weaknesses in dynamic correctness. Our contributions include a dataset of examples across three categories of spreadsheet tasks, an open source harness and agentic evaluation framework, and a characterization of existing frontier models' performance on our benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。