arXiv:2608.28155cs.CL2026-08

金融推理新基准FinExam-10K,验证检索增强模型在真实考试中的表现。

FinExam-10K: When Retrieval Helps Financial Reasoning?

论文配图:FinExam-10K: When Retrieval Helps Financial Reasoning?
图 1 · 摘自论文原文
  • 构建覆盖CFA与FRM全阶段的10,198题大型金融考试数据集
  • 检索增强模型提升准确率仅0.4%,但能修复数百个错误答案
  • 基于公开数据训练的门控机制,仅对7.9%问题触发高级推理

专业金融考试要求模型结合领域知识、计算与判断能力,但现有基准未统一覆盖CFA和FRM全部层级。我们提出FinExam-10K,据知是该场景下最大的英语基准,包含10,198道专家标注题目,涵盖CFA Levels I-III与FRM Parts I-II。释放5,110题用于训练,保留5,088题用于季度维护排行榜。为区分覆盖率与上下文可推理性,报告10,198项全覆盖赛道和7,625项上下文完备推理赛道。17个模型中最高准确率达85.29%。在冻结的难样本集上,全覆盖赛道最佳得分为34.68%,上下文完备372题中为54.57%。所有模型均在47个上下文完备问题上失败。Function-RAG与FunctionGraph-RAG修复数百错误,但也推翻大量正确答案,净收益微弱或为负。仅用公开数据训练的门控模型根据问题与初始回答决定是否调用FunctionGraph-RAG。在5,088个保留题上,门控触发率为7.9%,准确率从70.83%提升至71.23%(p = .0446)。

原文摘要 · Abstract (English)

Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for this setting, with 10,198 expert-reannotated questions spanning CFA Levels I-III and FRM Parts I-II. We release 5,110 questions and sequester 5,088 for a quarterly maintained leaderboard. To separate coverage from local answerability, we report a 10,198-item Full-Coverage Track and a 7,625-item Context-Complete Reasoning Track, which is the primary basis for claims about reasoning from the supplied record. Across 17 models, the best accuracy is 85.29% overall. On the frozen Hard band, the best score is 34.68% on the Full-Coverage Track and 54.57% on the 372 context-complete items. All 17 models share 47 context-complete failures. Function-RAG and FunctionGraph-RAG rescue hundreds of errors but also overturn many correct answers, producing little or negative net gain. A gate trained only on public data decides from the question and initial response when FunctionGraph-RAG should run. On the 5,088 held-out items, the gate invokes FunctionGraph-RAG for 7.9% of questions and improves accuracy from 70.83% to 71.23% (p = .0446).

金融推理检索增强基准测试门控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。