arXiv:2509.04468cs.CLcs.AI2025-09被引 4

用CFA真题评估大模型金融推理能力,发现专注推理的模型更优。

Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study

  • 用CFA三级1560道真题测试大模型,覆盖真实金融分析场景
  • 引入检索增强生成(RAG)系统,显著提升复杂问题解答准确率
  • 发现知识缺失是主要失败原因,适合金融领域模型选型参考

大语言模型在金融应用中潜力巨大,但专业领域的系统性评估仍显不足。本研究首次基于全球最严格的金融认证CFA Levels I-III的1,560道多选题,对前沿LLMs进行全面评估。对比了多模态强算力、专精推理高精度、轻量高效三类模型,在零样本提示与创新的检索增强生成(RAG)管道下进行测试。RAG通过分层知识组织与结构化查询生成,实现精准领域知识检索,显著提升专业金融认证评估中的推理准确性。结果显示,推理导向模型在零样本设置下持续领先,而RAG管道在复杂场景中带来显著改进。全面错误分析表明,知识缺口是主要失败模式,文本可读性影响极小。研究为金融领域大模型部署提供实证指导,助力模型选型与成本-性能优化。

原文摘要 · Abstract (English)

The rapid advancement of large language models presents significant opportunities for financial applications, yet systematic evaluation in specialized financial contexts remains limited. This study presents the first comprehensive evaluation of state-of-the-art LLMs using 1,560 multiple-choice questions from official mock exams across Levels I-III of CFA, most rigorous professional certifications globally that mirror real-world financial analysis complexity. We compare models distinguished by core design priorities: multi-modal and computationally powerful, reasoning-specialized and highly accurate, and lightweight efficiency-optimized. We assess models under zero-shot prompting and through a novel Retrieval-Augmented Generation pipeline that integrates official CFA curriculum content. The RAG system achieves precise domain-specific knowledge retrieval through hierarchical knowledge organization and structured query generation, significantly enhancing reasoning accuracy in professional financial certification evaluation. Results reveal that reasoning-oriented models consistently outperform others in zero-shot settings, while the RAG pipeline provides substantial improvements particularly for complex scenarios. Comprehensive error analysis identifies knowledge gaps as the primary failure mode, with minimal impact from text readability. These findings provide actionable insights for LLM deployment in finance, offering practitioners evidence-based guidance for model selection and cost-performance optimization.

金融推理大模型评估RAGCFA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。