arXiv:2608.04200q-fin.MFcs.LG2026-08

用QLoRA提升金融情感模型,发现准确率不等于可交易信号

From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models

论文配图:From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models
图 1 · 摘自论文原文
  • 采用QLoRA微调大模型,显著提升金融文本分类效果
  • 最佳模型在测试集上准确率达88.40%,但预测收益能力很弱
  • 适合关注金融AI模型经济有效性的研究者和量化从业者

金融情感分类器通常以人工标注为评价标准,但语言性能强未必带来经济上有用的收益预测。本研究通过两个实验分离这两个问题:首先,在五个金融文本数据集基础上构建统一三分类基准,对比TF-IDF朴素贝叶斯、现成的FinBERT与Financial-RoBERTa编码器、零样本Qwen2.5-7B,以及经QLoRA微调的Qwen2.5-7B、LLaMA3-8B和Mistral-7B模型。Mistral-7B在测试集上达到最高准确率(0.8840)和宏平均F1(0.8771),QLoRA将Qwen2.5的宏平均F1从0.7274提升至0.8615。使用逆频率类别加权损失未改善Qwen2.5表现。其次,在2019年独立的Benzinga数据集(含10,637条新闻标题,13,115个标题-股票观测)上评估经济有效性,固定S&P 100标的。模型概率转为连续情感分数,按股票与信号日聚合,并与后续1、2、3、5日收益率对齐。所有七个下游模型在一日期 horizon 上均产生正但微小的平均秩信息系数(最大为0.0143,来自FinBERT),28组模型-周期组合在经过Newey-West修正与错误发现率校正后均不显著。投资组合结果亦未显示最优分类器有稳健优势。结果表明,QLoRA有效提升金融情感建模能力,同时揭示分类准确率与可交易横截面信号间存在明显鸿沟。

原文摘要 · Abstract (English)

Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10,637 unique headlines and 13,115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.

金融AIQLoRA情绪分析量化投资

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。