23个大模型在CFA三级考试中表现亮眼,最高准确率达79.1%。
Advanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III
- 用链式思考与自发现提示评估大模型金融推理能力
- o4-mini模型在论文评分中达到79.1%准确率
- 适合金融从业者评估AI模型选型参考
随着金融机构越来越多采用大语言模型(LLMs),进行领域专用的严格评估对负责任部署至关重要。本文提出一个综合性基准,评估23个前沿大模型在特许金融分析师(CFA)Level III考试中的表现——该考试是高级金融推理的黄金标准。我们采用多种提示策略(包括链式思考和自发现),评估多选题(MCQs)和论述题回答。评估结果显示,领先模型展现出强大能力,如o4-mini模型在新修订的更严格论述题评分标准下达到79.1%的综合得分,Gemini 2.5 Flash为77.3%。这些结果表明大模型在高风险金融应用中已取得显著进展。研究为从业者提供关键模型选型指导,同时指出在成本效益部署及专业基准性能解读方面仍存挑战。
原文摘要 · Abstract (English)
As financial institutions increasingly adopt Large Language Models (LLMs), rigorous domain-specific evaluation becomes critical for responsible deployment. This paper presents a comprehensive benchmark evaluating 23 state-of-the-art LLMs on the Chartered Financial Analyst (CFA) Level III exam - the gold standard for advanced financial reasoning. We assess both multiple-choice questions (MCQs) and essay-style responses using multiple prompting strategies including Chain-of-Thought and Self-Discover. Our evaluation reveals that leading models demonstrate strong capabilities, with composite scores such as 79.1% (o4-mini) and 77.3% (Gemini 2.5 Flash) on CFA Level III. These results, achieved under a revised, stricter essay grading methodology, indicate significant progress in LLM capabilities for high-stakes financial applications. Our findings provide crucial guidance for practitioners on model selection and highlight remaining challenges in cost-effective deployment and the need for nuanced interpretation of performance against professional benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。