首个面向金融领域AI代理的综合性评估基准,测试其真实业务能力。
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
- 构建涵盖7大金融子领域的407个任务,分三层场景深度评测
- 最佳模型ChatGPT零样本准确率仅48.9%,远低于专家水平
- 发现5类常见失败模式,为未来研究指明方向
AI代理的快速发展为自动化复杂任务带来新机遇,但其在金融领域的多步骤、多工具协作能力仍待探索。本文提出FinGAIA,一个面向真实金融场景的端到端评估基准,包含407个精心设计的任务,覆盖证券、基金、银行、保险、期货、信托和资产管理七大子领域,按基础业务分析、资产决策支持、战略风险管理三个层级组织。我们在零样本条件下评估了10个主流AI代理,表现最好的ChatGPT整体准确率为48.9%,虽优于非专业人士,但仍比金融专家低超过35个百分点。错误分析揭示出五类典型失败模式:跨模态对齐不足、金融术语偏差、操作流程意识缺失等。该工作首次提供贴近金融实际的代理评估基准,旨在客观评估并推动该关键领域的进展。部分数据已开源于https://github.com/SUFE-AIFLM-Lab/FinGAIA。
原文摘要 · Abstract (English)
The booming development of AI agents presents unprecedented opportunities for automating complex tasks across various domains. However, their multi-step, multi-tool collaboration capabilities in the financial sector remain underexplored. This paper introduces FinGAIA, an end-to-end benchmark designed to evaluate the practical abilities of AI agents in the financial domain. FinGAIA comprises 407 meticulously crafted tasks, spanning seven major financial sub-domains: securities, funds, banking, insurance, futures, trusts, and asset management. These tasks are organized into three hierarchical levels of scenario depth: basic business analysis, asset decision support, and strategic risk management. We evaluated 10 mainstream AI agents in a zero-shot setting. The best-performing agent, ChatGPT, achieved an overall accuracy of 48.9\%, which, while superior to non-professionals, still lags financial experts by over 35 percentage points. Error analysis has revealed five recurring failure patterns: Cross-modal Alignment Deficiency, Financial Terminological Bias, Operational Process Awareness Barrier, among others. These patterns point to crucial directions for future research. Our work provides the first agent benchmark closely related to the financial domain, aiming to objectively assess and promote the development of agents in this crucial field. Partial data is available at https://github.com/SUFE-AIFLM-Lab/FinGAIA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。