arXiv:2502.10990cs.CLcs.IR2025-02EMNLP被引 26

为金融领域打造文本嵌入评估基准,揭示现有模型的局限性。

FinMTEB: Finance Massive Text Embedding Benchmark

  • 构建64个金融数据集,覆盖7类任务和中英文文本类型。
  • 通用模型在金融任务上表现不佳,领域适配模型显著更优。
  • 简单词袋法竟胜过复杂嵌入模型,暴露密集表示缺陷。

嵌入模型在各类自然语言处理应用中起关键作用。近年来大语言模型的发展进一步提升了嵌入模型性能。尽管这些模型常在通用数据集上评估,但实际应用需领域特定评价。本文提出金融领域专用的文本嵌入基准FinMTEB,作为MTEB的金融对应版本。FinMTEB包含64个金融领域嵌入数据集,涵盖7项任务,覆盖中英文多种文本类型,如财经新闻、企业年报、ESG报告、监管文件和财报电话会议记录。我们还开发了金融适配模型Fin-E5,采用基于角色的数据合成方法以覆盖多样化的金融嵌入任务。对15个嵌入模型(包括Fin-E5)的全面评估显示:(1)通用基准表现与金融任务相关性低;(2)领域适配模型始终优于通用模型;(3)出人意料的是,简单的词袋法(BoW)在金融语义文本相似度(STS)任务上优于复杂密集嵌入,凸显当前密集嵌入技术的局限性。本工作建立了金融NLP应用的稳健评估框架,并为开发领域专用嵌入模型提供关键洞见。

原文摘要 · Abstract (English)

Embedding models play a crucial role in representing and retrieving information across various NLP applications. Recent advances in large language models (LLMs) have further enhanced the performance of embedding models. While these models are often benchmarked on general-purpose datasets, real-world applications demand domain-specific evaluation. In this work, we introduce the Finance Massive Text Embedding Benchmark (FinMTEB), a specialized counterpart to MTEB designed for the financial domain. FinMTEB comprises 64 financial domain-specific embedding datasets across 7 tasks that cover diverse textual types in both Chinese and English, such as financial news articles, corporate annual reports, ESG reports, regulatory filings, and earnings call transcripts. We also develop a finance-adapted model, Fin-E5, using a persona-based data synthetic method to cover diverse financial embedding tasks for training. Through extensive evaluation of 15 embedding models, including Fin-E5, we show three key findings: (1) performance on general-purpose benchmarks shows limited correlation with financial domain tasks; (2) domain-adapted models consistently outperform their general-purpose counterparts; and (3) surprisingly, a simple Bag-of-Words (BoW) approach outperforms sophisticated dense embeddings in financial Semantic Textual Similarity (STS) tasks, underscoring current limitations in dense embedding techniques. Our work establishes a robust evaluation framework for financial NLP applications and provides crucial insights for developing domain-specific embedding models.

文本嵌入金融NLP评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。