arXiv:2512.08270cs.AIcs.CL2025-12被引 2

顶尖推理模型在CFA三级别考试中全部通过,表现远超以往语言模型。

Reasoning Models Ace the CFA Exams

  • 用最新推理模型测试980道CFA模拟题,覆盖三个级别。
  • Gemini 3.0 Pro在Level I达97.6%正确率,创纪录;GPT-5在Level II达94.3%。
  • 适合关注AI在金融专业能力评估中表现的研究者与从业者。

先前研究显示大型语言模型(LLMs)在特许金融分析师(CFA)考试中表现不佳。然而,近期推理模型在多个学科的研究生级学术与专业考试中取得优异成绩。本文评估了最先进的推理模型在一套包含980道题的CFA模拟试卷上的表现,涵盖三个Level I、两个Level II和三个Level III考试。采用与前期研究相同的及格标准,发现大多数模型均能通过所有三个级别。按综合表现排序,通过模型依次为Gemini 3.0 Pro、Gemini 2.5 Pro、GPT-5、Grok 4、Claude Opus 4.1和DeepSeek-V3.1。具体而言,Gemini 3.0 Pro在Level I达到97.6%的记录分数;GPT-5在Level II以94.3%领先;在Level III,Gemini 2.5 Pro在选择题上获得最高分86.4%,而Gemini 3.0 Pro在主观题上达到92.0%。

原文摘要 · Abstract (English)

Previous research has reported that large language models (LLMs) demonstrate poor performance on the Chartered Financial Analyst (CFA) exams. However, recent reasoning models have achieved strong results on graduate-level academic and professional examinations across various disciplines. In this paper, we evaluate state-of-the-art reasoning models on a set of mock CFA exams consisting of 980 questions across three Level I exams, two Level II exams, and three Level III exams. Using the same pass/fail criteria from prior studies, we find that most models clear all three levels. The models that pass, ordered by overall performance, are Gemini 3.0 Pro, Gemini 2.5 Pro, GPT-5, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1. Specifically, Gemini 3.0 Pro achieves a record score of 97.6% on Level I. Performance is also strong on Level II, led by GPT-5 at 94.3%. On Level III, Gemini 2.5 Pro attains the highest score with 86.4% on multiple-choice questions while Gemini 3.0 Pro achieves 92.0% on constructed-response questions.

CFA考试推理模型金融AI大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。