arXiv:2412.09859cs.LGcs.CL2024-12被引 2

用真实与合成数据提升金融文本情感分析效果

Financial Sentiment Analysis: Leveraging Actual and Synthetic Data for Supervised Fine-tuning

  • 融合真实与合成数据微调模型,增强金融语境理解
  • 在50%和100%一致标准下,准确率与F1值均提升
  • 适合金融量化、投资决策等需要情绪信号的场景

有效市场假说(EMH)强调了金融新闻对股价变动的核心作用。金融新闻以公司公告、新闻标题等形式存在,其内容可通过情感分析转化为投资洞察。通用语言模型在金融领域表现不足,且用于微调的标注数据稀缺,现有金融情感分析模型也未能充分利用长上下文。本文提出结合真实与合成数据的方法,构建BertNSP-finance以拼接短金融句形成长句,采用finbert-lc进行情感判断。实验结果表明,在金融短语银行数据集上,于50%和100%一致标准下,准确率与F1值均获得显著提升。

原文摘要 · Abstract (English)

The Efficient Market Hypothesis (EMH) highlights the essence of financial news in stock price movement. Financial news comes in the form of corporate announcements, news titles, and other forms of digital text. The generation of insights from financial news can be done with sentiment analysis. General-purpose language models are too general for sentiment analysis in finance. Curated labeled data for fine-tuning general-purpose language models are scare, and existing fine-tuned models for sentiment analysis in finance do not capture the maximum context width. We hypothesize that using actual and synthetic data can improve performance. We introduce BertNSP-finance to concatenate shorter financial sentences into longer financial sentences, and finbert-lc to determine sentiment from digital text. The results show improved performance on the accuracy and the f1 score for the financial phrasebank data with $50\%$ and $100\%$ agreement levels.

金融文本情感分析合成数据微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。