首个覆盖金融全流程的LLM评估基准,测试复杂场景下的推理能力。
FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
- 构建三模块框架:数据生成、任务设计与统一评估
- 复杂任务准确率仅40%,多步推理错误会累积放大
- 适合研究金融AI的开发者和金融科技从业者
金融任务对全球经济稳定至关重要,但执行中面临人工密集、容错率低、数据分散和工具限制等问题。尽管大语言模型在自然语言处理中表现优异,且具备通过推理和上下文理解自动化流程的潜力,但现有金融领域评估基准缺乏领域特异性数据、任务设计简单且评价体系不完整。为此,本文提出FinMaster,一个全面的金融基准,系统评估大模型在金融素养、会计、审计和咨询等领域的综合能力。该基准包含三个核心模块:i)FinSim,构建可生成合成隐私合规财务数据的模拟器,以复现市场动态;ii)FinSuite,提供涵盖183种不同类型与难度的金融核心任务;iii)FinEval,开发统一评估接口。对先进大模型的广泛实验揭示了关键能力差距:在需要多步推理的复杂场景中,准确率从基础任务的90%以上骤降至40%。计算误差具有传播性,单一指标计算准确率为58%,在多指标场景中降至37%。据我们所知,FinMaster是首个覆盖全流程金融工作流且任务挑战性强的基准。期望其能弥合研究与产业间的鸿沟,推动大模型在真实金融实践中的应用,提升效率与准确性。
原文摘要 · Abstract (English)
Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, and tool limitations. Although large language models (LLMs) have succeeded in various natural language processing tasks and have shown potential in automating workflows through reasoning and contextual understanding, current benchmarks for evaluating LLMs in finance lack sufficient domain-specific data, have simplistic task design, and incomplete evaluation frameworks. To address these gaps, this article presents FinMaster, a comprehensive financial benchmark designed to systematically assess the capabilities of LLM in financial literacy, accounting, auditing, and consulting. Specifically, FinMaster comprises three main modules: i) FinSim, which builds simulators that generate synthetic, privacy-compliant financial data for companies to replicate market dynamics; ii) FinSuite, which provides tasks in core financial domains, spanning 183 tasks of various types and difficulty levels; and iii) FinEval, which develops a unified interface for evaluation. Extensive experiments over state-of-the-art LLMs reveal critical capability gaps in financial reasoning, with accuracy dropping from over 90% on basic tasks to merely 40% on complex scenarios requiring multi-step reasoning. This degradation exhibits the propagation of computational errors, where single-metric calculations initially demonstrating 58% accuracy decreased to 37% in multimetric scenarios. To the best of our knowledge, FinMaster is the first benchmark that covers full-pipeline financial workflows with challenging tasks. We hope that FinMaster can bridge the gap between research and industry practitioners, driving the adoption of LLMs in real-world financial practices to enhance efficiency and accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。