arXiv:2605.15482cs.CL2026-05被引 2

构建金融领域分层评估基准,测试大模型从基础到专家级的财务能力。

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models

论文配图:FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models
图 1 · 摘自论文原文
  • 按专业认证难度分级设计8个子任务,覆盖CFA/CMT等金融考试层级
  • 包含3993道题,可评估模型在计算、判断与开放回答中的表现退化
  • 支持多类型答案自动评分,适合金融大模型专业能力验证

大型语言模型在金融分析、投资决策、风险管理和合规等领域应用日益广泛,但其金融专业能力的稳健评估仍不完善。现有公开基准如FinQA、ConvFinQA和TAT-QA主要聚焦财务报告问答,缺乏专业难度层级。更广泛的资源如FinanceBench、PIXIU、FinBen和FLaME虽扩展了任务范围,但仍未解决从基础认知向专家推理过渡的评估问题。本文提出FINESSE-Bench,一个包含8个专业化评测子集的分层基准,共3,993道题目,涵盖模拟CFA一级至三级、CMT二级、CFTe一级等专业认证题型,以及实际交易任务和俄语金融奥赛数据。该设计支持对领域广度、难度上升下的性能下降、计算任务求解能力及专精领域行为的评估。我们还提出统一评估协议,涵盖选择题、数值答案与简答,并基于大模型作为裁判的自动评分机制实现自由回答评分。FINESSE-Bench旨在补充现有金融基准,推动对大模型金融专业能力的实质性评估。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being applied to financial analysis, reporting, investment decision support, risk management, compliance, and professional training. However, robust evaluation of their domain competence in finance remains incomplete. Widely used open benchmarks such as FinQA, ConvFinQA, and TAT-QA have played an important role in advancing financial question answering and numerical reasoning, but they focus primarily on question answering over financial reports and do not provide an explicit hierarchy of professional difficulty. Broader resources, including FinanceBench, PIXIU, FinBen, and FLaME, expand the coverage of financial tasks, yet the problem of evaluating the transition from foundational knowledge to expert-level financial reasoning remains open. In this work, we present FINESSE-Bench, a suite of eight specialized benchmarks comprising 3,993 questions for hierarchical evaluation of financial competencies in LLMs. FINESSE-Bench combines exam-oriented datasets inspired by professional certifications (CFA-like Levels 1-3, CMT-like Level 2, and CFTe-like Level 1), applied trading task collections, and a Russian-language olympiad benchmark. This design enables evaluation of domain breadth, performance degradation as difficulty increases, the ability to solve computational tasks, and model behavior in specialized financial domains. We also describe a unified evaluation protocol covering multiple-choice questions, numerical answers, and short open-ended responses, together with an automated scoring scheme for freeform answers based on the LLM-as-judge paradigm. FINESSE-Bench is intended both as a complement to existing open financial benchmarks and as a tool for more substantive evaluation of professionally relevant financial competencies in large language models.

金融AI模型评估分层测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。