arXiv:2412.19140cs.CLcs.AI2024-12被引 2

用自洽提示修正提升金融实体情感分析效果,构建最大中英文数据集。

SILC-EFSA: Self-aware In-context Learning Correction for Entity-level Financial Sentiment Analysis

  • 两阶段框架:先生成伪标签,再用图神经网络检索修正。
  • 在新构建的中英文数据集上达当前最佳性能。
  • 适合金融舆情监控与量化研究者使用。

近年来,金融领域细粒度情感分析受到广泛关注,但实体级数据集稀缺仍是关键挑战。为此,我们构建了迄今最大的中英文金融实体级情感分析数据集。基于此,提出新型两阶段情感分析方法SILC(Self-aware In-context Learning Correction)。第一阶段通过微调大语言模型生成特定任务的伪标签数据;第二阶段利用基于图神经网络的示例检索器,结合伪标签数据训练修正模型。该两阶段策略在新构建的数据集上取得当前最优性能,推动金融情感分析发展。案例研究显示,本方法在加密货币市场监测中具备更强实用性。数据与代码已开源:https://github.com/NLP-Bin/SILC-EFSA。

原文摘要 · Abstract (English)

In recent years, fine-grained sentiment analysis in finance has gained significant attention, but the scarcity of entity-level datasets remains a key challenge. To address this, we have constructed the largest English and Chinese financial entity-level sentiment analysis datasets to date. Building on this foundation, we propose a novel two-stage sentiment analysis approach called Self-aware In-context Learning Correction (SILC). The first stage involves fine-tuning a base large language model to generate pseudo-labeled data specific to our task. In the second stage, we train a correction model using a GNN-based example retriever, which is informed by the pseudo-labeled data. This two-stage strategy has allowed us to achieve state-of-the-art performance on the newly constructed datasets, advancing the field of financial sentiment analysis. In a case study, we demonstrate the enhanced practical utility of our data and methods in monitoring the cryptocurrency market. Our datasets and code are available at https://github.com/NLP-Bin/SILC-EFSA.

情感分析金融大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。