arXiv:2506.15846cs.CLcs.AI2025-06被引 4

首个金融语言模型评估基准,揭示大模型在金融任务中的真实潜力

Finance Language Model Evaluation (FLaME)

  • 构建首个全面的金融语言模型评测框架FLaME
  • 23个基础模型在20项金融任务上实证表现优于传统认知
  • 开源全部数据与结果,推动金融NLP研究透明化

语言模型在通用自然语言处理任务中表现出色,但在金融领域等高度专业化、知识密集型任务中的实际效能仍难以准确评估,原因在于现有评估方法存在重大缺陷,导致对语言模型在常见金融NLP(FinNLP)任务上性能的低估。为展示语言模型在这些任务中的真实潜力,我们提出了首个金融语言模型评估的综合基准测试套件FLaME。本研究首次系统性地对比了普通语言模型与‘推理增强型’语言模型,在23个主流基础模型和20个核心金融NLP任务上的表现。我们开源了整个框架、所有数据及实验结果,以促进金融语言模型研究的可复现性与进展。

原文摘要 · Abstract (English)

Language Models (LMs) have demonstrated impressive capabilities with core Natural Language Processing (NLP) tasks. The effectiveness of LMs for highly specialized knowledge-intensive tasks in finance remains difficult to assess due to major gaps in the methodologies of existing evaluation frameworks, which have caused an erroneous belief in a far lower bound of LMs' performance on common Finance NLP (FinNLP) tasks. To demonstrate the potential of LMs for these FinNLP tasks, we present the first holistic benchmarking suite for Financial Language Model Evaluation (FLaME). We are the first research paper to comprehensively study LMs against 'reasoning-reinforced' LMs, with an empirical study of 23 foundation LMs over 20 core NLP tasks in finance. We open-source our framework software along with all data and results.

金融NLP模型评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。