为金融领域智能体设计测评基准,验证其在真实操作中的可靠性。
FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance

- 构建8维评分体系,聚焦金融场景下的准确、可验证与合规性。
- 在251个专家标注问题上测试,通用智能体多数不达标。
- 适合金融从业者、AI研发者用于评估企业级智能体性能。
大型语言模型的进展推动了智能体系统在企业财务中的部署。现有基准多关注通用能力、指令遵循或安全性,却很少针对实际财务工作流程进行评估。财务人员要求智能体不仅提供事实准确且有依据的信息,还需确保信息可验证,并严格遵守财务领域的规则与约束。我们提出FORCE-Bench,包含251个专家标注的查询任务,采用基于评分标准的评估框架,从准确性、引用、清晰度、深度、可溯源性、时效性、相关性和结构共八个维度评估智能体表现。评估涵盖三类任务:财务义务研究(查询ERP系统中应收应付数据)、财务实体绩效研究(回答基于公开文件和市场数据的时间限定问题)以及商业简报生成(整合多源公司情报报告)。为模拟真实部署环境,我们在常见工具访问权限和延迟限制条件下,评估自研金融智能体及通用智能体系统。结果表明,通用智能体在操作约束下无法稳定满足金融领域质量要求,而专为Microsoft 365 Copilot设计的金融智能体表现更可靠。我们开源数据集、评分标准、评估工具与分析代码,支持可复现对比并适配其他企业财务场景。
原文摘要 · Abstract (English)
Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring general capabilities, instruction following, or safety, but few directly address the operational finance workflows that agentic systems are now being deployed to automate. Finance professionals require agents to not only provide factually sound and properly grounded information, but also ensure that this information is verifiable and consistently adheres to rules and constraints of the operational finance domain. We introduce FORCE-Bench, which contains 251 expert-annotated queries and evaluates responses using a rubric-based framework calibrated to the requirements of the operational finance domain, across eight dimensions: accuracy, citations, clarity, depth, groundedness, recency, relevance, and structure. FORCE-Bench assesses agentic systems on three task types: financial obligation research (querying ERP systems for accounts receivable and payable data), financial entity performance research (answering time-bound questions from public filings and market data), and business brief generation (synthesising multi-source company intelligence reports). To reflect real deployment conditions, we evaluate our purpose-built agent, as well as the general-purpose agentic systems, under common tool access and latency-bounded settings. Results show that general-purpose agentic systems do not consistently meet finance-domain quality requirements under operational constraints, while the purpose-built Finance Agent for Microsoft 365 Copilot is more reliable across dimensions. We release the dataset, rubrics, harness, and analysis code as open-source to support reproducible comparison and adaptation to other enterprise finance environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。