构建金融数值推理新基准,更可信、更全面、更具挑战性。
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
- 引入真实金融问题与Python解法,提升评估可信度。
- 覆盖67.8%金融概念,3133个函数提升模型表现至91.6%准确率。
- 设置238道高难度题,适合研究金融推理模型的开发者参考。
我们提出FinanceReasoning,一个用于评估大语言模型在金融数值推理任务中推理能力的新基准。相比现有基准,本工作实现三大突破:(1)可信性:从四个公开数据集更新15.6%题目,新增908道带详细Python解法的问题,并严格优化评估标准,实现对模型推理进步的精准评估;(2)全面性:覆盖67.8%的金融概念与公式,构建3,133个以Python格式呈现的函数,显著提升模型能力(如GPT-4o准确率从83.2%提升至91.6%);(3)挑战性:要求模型在238道难题中综合运用多个金融公式进行精确推理。当前最佳模型(OpenAI o1 with PoT)达到89.1%准确率,仍面临数值精度挑战。我们发现结合Reasoner与Programmer模型可有效提升性能(如DeepSeek-R1从83.2%升至87.8%)。该工作为领域特定复杂推理任务的评估与改进提供新方向。
原文摘要 · Abstract (English)
We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compared to existing benchmarks, our work provides three key advancements. (1) Credibility: We update 15.6% of the questions from four public datasets, annotating 908 new questions with detailed Python solutions and rigorously refining evaluation standards. This enables an accurate assessment of the reasoning improvements of LRMs. (2) Comprehensiveness: FinanceReasoning covers 67.8% of financial concepts and formulas, significantly surpassing existing datasets. Additionally, we construct 3,133 Python-formatted functions, which enhances LRMs' financial reasoning capabilities through refined knowledge (e.g., 83.2% $\rightarrow$ 91.6% for GPT-4o). (3) Challenge: Models are required to apply multiple financial formulas for precise numerical reasoning on 238 Hard problems. The best-performing model (i.e., OpenAI o1 with PoT) achieves 89.1% accuracy, yet LRMs still face challenges in numerical precision. We demonstrate that combining Reasoner and Programmer models can effectively enhance LRMs' performance (e.g., 83.2% $\rightarrow$ 87.8% for DeepSeek-R1). Our work paves the way for future research on evaluating and improving LRMs in domain-specific complex reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。