arXiv:2512.13168cs.AIcs.CE2025-12ACL被引 7

评测AI在真实企业财务流程中的表现,发现顶级模型仍难搞定复杂任务。

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows

  • 基于真实企业邮件和表格历史,构建含384个任务的财务工作流数据集。
  • 172个复合任务覆盖预算、交易等场景,涉及2700万单元格与多类文档。
  • 顶级AI模型平均耗时16.8分钟仅通过38.4%任务,暴露实际应用短板。

我们提出FinWorkBench(简称Finch),用于评估AI代理在真实企业级财务与会计工作流中的表现,这些流程涵盖数据录入、结构化、格式化、网页搜索、跨文件检索、计算、建模、验证、翻译、可视化及报告生成。数据源自埃克森(Enron)(15,000份文件、50万封邮件)及其他金融机构的真实工作空间,时间跨度为2000–2025年,保留了多模态资料(如表格、图表)在预算、交易、资产管理与运营等领域的原始混乱状态。通过结合大模型辅助挖掘与专家标注的流程构建方法:(1)从真实邮件与电子表格版本历史中提取工作流并经专家验证;(2)投入超700小时专家标注。最终获得172个复合工作流,共384项任务,涉及1,710个电子表格(2700万单元格)以及PDF等其他资料,完整呈现真实企业工作流的长期性、知识密集性与协作特性。我们对前沿AI系统(包括GPT-5.1、Claude Sonnet 4.5、Claude Opus 4.5、Gemini 3 Pro、Grok 4、Qwen 3 Max)进行了人工与自动评估。人工评测显示,GPT-5.1 Pro平均每工作流耗时16.8分钟,但仅通过38.4%的任务。深入案例研究揭示了企业工作流对AI代理带来的多重挑战。

原文摘要 · Abstract (English)

We introduce FinWorkBench (a.k.a. Finch) for evaluating AI agents on real-world, enterprise-grade finance and accounting workflows that interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting. Finch is sourced from authentic enterprise workspaces from Enron (15,000 files and 500,000 emails) and other financial institutions, covering the period 2000--2025 and preserving the in-the-wild messiness of multimodal artifacts such as tables and charts across diverse domains including budgeting, trading, asset management, and operational management. We propose a workflow construction process that combines LLM-assisted mining of workflows from authentic enterprise environments with expert annotation: (1) LLM-assisted, expert-verified derivation of workflows from real-world email threads and spreadsheet version histories, and (2) meticulous annotation requiring over 700 hours of expert effort. This yields 172 composite workflows with 384 tasks, involving 1,710 spreadsheets with 27 million cells, along with PDFs and other artifacts, capturing the intrinsically messy, long-horizon, knowledge-intensive, and collaborative nature of real-world enterprise work. We conduct both human and automated evaluations of frontier AI systems, including GPT-5.1, Claude Sonnet 4.5, Claude Opus 4.5, Gemini 3 Pro, Grok 4, and Qwen 3 Max. Under human evaluation, GPT-5.1 Pro spends an average of 16.8 minutes per workflow yet passes only 38.4% of workflows. Comprehensive case studies further surface the challenges that real-world enterprise workflows pose for AI agents.

财务AI工作流评测企业应用大模型测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。