小模型经微调后在金融新闻情感分析上超越GPT-3.5和GPT-4。
Optimizing Performance: How Compact Models Match or Exceed GPT's Classification Capabilities through Fine-Tuning
- 用自建市场评分数据库微调小模型,避免人为偏见。
- 微调后的FinBERT等模型在零样本学习中超越GPT-3.5/4。
- 模型行为相似性暗示非独立性,适合金融文本分类研究者。
本文证明,经过微调的非生成式小型模型(如FinBERT和FinDRoBERTa)在金融新闻情感分析的零样本学习设置下,可超越GPT-3.5和GPT-4的表现。这些微调模型在基于Bloomberg每日财经摘要的任务中,与微调后的GPT-3.5表现相当。为实现对比,研究构建了一个新型数据库,通过系统识别新闻提及公司并分析其股价涨跌(上涨、下跌或中性),自动赋予每条新闻市场得分,消除人工解读偏差。此外,研究发现康多塞陪审团定理的前提不成立,表明微调后的小型模型与微调后的GPT模型存在行为相似性,非相互独立。最后,所有微调模型已公开发布于HuggingFace,供后续金融情感分析与文本分类研究使用。
原文摘要 · Abstract (English)
In this paper, we demonstrate that non-generative, small-sized models such as FinBERT and FinDRoBERTa, when fine-tuned, can outperform GPT-3.5 and GPT-4 models in zero-shot learning settings in sentiment analysis for financial news. These fine-tuned models show comparable results to GPT-3.5 when it is fine-tuned on the task of determining market sentiment from daily financial news summaries sourced from Bloomberg. To fine-tune and compare these models, we created a novel database, which assigns a market score to each piece of news without human interpretation bias, systematically identifying the mentioned companies and analyzing whether their stocks have gone up, down, or remained neutral. Furthermore, the paper shows that the assumptions of Condorcet's Jury Theorem do not hold suggesting that fine-tuned small models are not independent of the fine-tuned GPT models, indicating behavioural similarities. Lastly, the resulted fine-tuned models are made publicly available on HuggingFace, providing a resource for further research in financial sentiment analysis and text classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。