arXiv:2411.07142cs.CL2024-11EMNLP被引 13

金融文本专用嵌入模型,比通用模型召回率高23.6个百分点。

Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt

  • 基于1430万条金融问答对微调,专攻金融术语与长尾表达
  • 在测试集上召回率62.8%,超越最强通用模型23.6个百分点
  • 特别擅长处理含公司名、时间点的复杂金融查询,适合量化研究者

金融文档充斥着专业术语、晦涩行话和奇特缩写,对通用文本嵌入构成挑战。但目前鲜有金融专用嵌入模型发表,部分原因在于缺乏公开数据集与基准。我们提出BAM嵌入,基于精心构建的1430万条查询-段落对进行微调。实验证明领域特化训练的优势:BAM在保留测试集上的召回率(Recall@1)达62.8%,而最佳通用嵌入(OpenAI)仅为39.2%。此外,BAM在FinanceBench上将问答准确率提升8%,对包含公司名、日期及前瞻性表述的金融特异性查询表现出更强敏感性。为推动后续研究,本文详细描述方法,量化硬负样本挖掘与数据规模的重要性。

原文摘要 · Abstract (English)

Financial documents are filled with specialized terminology, arcane jargon, and curious acronyms that pose challenges for general-purpose text embeddings. Yet, few text embeddings specialized for finance have been reported in the literature, perhaps in part due to a lack of public datasets and benchmarks. We present BAM embeddings, a set of text embeddings finetuned on a carefully constructed dataset of 14.3M query-passage pairs. Demonstrating the benefits of domain-specific training, BAM embeddings achieve Recall@1 of 62.8% on a held-out test set, vs. only 39.2% for the best general-purpose text embedding from OpenAI. Further, BAM embeddings increase question answering accuracy by 8% on FinanceBench and show increased sensitivity to the finance-specific elements that are found in detailed, forward-looking and company and date-specific queries. To support further research we describe our approach in detail, quantify the importance of hard negative mining and dataset scale.

金融嵌入文本检索领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。