arXiv:2504.16188cs.CLcs.AI2025-04NAACL被引 6

构建金融领域自然语言推理新数据集,揭示大模型在金融文本理解中的短板。

FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking

  • 构建跨多类型金融文本的推理数据集,减少表面关联干扰。
  • 包含21,304对样本,专家标注测试集达3,304条,基准模型最高F1仅78.62%。
  • 发现指令微调的金融大模型表现不佳,凸显领域适配性不足。

我们提出FinNLI,一个面向多种金融文本(如SEC文件、年报、财报电话会记录)的自然语言推理基准数据集。该数据集框架确保前提与假设对具有多样性,同时最小化表面相关性。FinNLI共包含21,304个样本对,其中3,304个为由金融专家标注的高质量测试集。评估显示,领域迁移显著降低通用NLI模型性能。预训练语言模型(PLMs)和大语言模型(LLMs)基线的最高宏观F1分别为74.57%和78.62%,表明该数据集难度高。令人意外的是,经过指令微调的金融专用大模型表现较差,提示其泛化能力有限。FinNLI揭示了当前大模型在金融推理任务中的不足,为后续研究指明方向。

原文摘要 · Abstract (English)

We introduce FinNLI, a benchmark dataset for Financial Natural Language Inference (FinNLI) across diverse financial texts like SEC Filings, Annual Reports, and Earnings Call transcripts. Our dataset framework ensures diverse premise-hypothesis pairs while minimizing spurious correlations. FinNLI comprises 21,304 pairs, including a high-quality test set of 3,304 instances annotated by finance experts. Evaluations show that domain shift significantly degrades general-domain NLI performance. The highest Macro F1 scores for pre-trained (PLMs) and large language models (LLMs) baselines are 74.57% and 78.62%, respectively, highlighting the dataset's difficulty. Surprisingly, instruction-tuned financial LLMs perform poorly, suggesting limited generalizability. FinNLI exposes weaknesses in current LLMs for financial reasoning, indicating room for improvement.

金融NLP自然语言推理大模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。