构建金融智能评估基准,测试大模型理论与实务能力
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
- 基于金融资格考题与业务场景设计双维度评测体系
- 涵盖3000个带参考答案的决策题与评分标准的开放题
- 适合金融AI研究者、模型开发者及风控系统评估者
我们提出FIRE,一个全面的金融智能与推理评估基准,用于评估大语言模型在金融理论知识和实际业务场景中的表现。针对理论能力,我们收集了来自广泛认可的金融资格考试的多样化试题,以检验模型对金融知识的深度理解与应用能力。为评估模型在真实金融任务中的实用性,我们设计了一套系统化评估矩阵,覆盖复杂金融领域的主要子领域与核心业务活动。基于该矩阵,我们构建了3000个金融情景问题,包括有明确答案的闭合式决策题和采用预设评分标准的开放式问题。我们在FIRE基准上对当前最先进的大模型(如XuanYuan 4.0,我们最新的金融领域模型)进行了全面评估,系统分析了现有大模型在金融应用中的能力边界。基准题集与评估代码已公开,以支持未来研究。
原文摘要 · Abstract (English)
We introduce FIRE, a comprehensive benchmark designed to evaluate both the theoretical financial knowledge of LLMs and their ability to handle practical business scenarios. For theoretical assessment, we curate a diverse set of examination questions drawn from widely recognized financial qualification exams, enabling evaluation of LLMs deep understanding and application of financial knowledge. In addition, to assess the practical value of LLMs in real-world financial tasks, we propose a systematic evaluation matrix that categorizes complex financial domains and ensures coverage of essential subdomains and business activities. Based on this evaluation matrix, we collect 3,000 financial scenario questions, consisting of closed-form decision questions with reference answers and open-ended questions evaluated by predefined rubrics. We conduct comprehensive evaluations of state-of-the-art LLMs on the FIRE benchmark, including XuanYuan 4.0, our latest financial-domain model, as a strong in-domain baseline. These results enable a systematic analysis of the capability boundaries of current LLMs in financial applications. We publicly release the benchmark questions and evaluation code to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。