对比大模型与词典法在模因股尾部风险预警中的表现
LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks

- 用大模型提取情绪极性、看涨度等多维情感信号
- 大模型信号与股价极端上涨有更强统计关联性
- 适合研究散户驱动市场波动的量化分析师
本文实证比较了基于词典和大语言模型(LLM)的情感分析方法,从社交媒体话语中提取高波动股票市场的相关信号。以r/WallStreetBets Reddit数据为对象,聚焦模因股(GME、AMC、NOK),构建时间对齐的情感指标,评估其与市场收益的关系,尤其关注收益分布上尾部极端正收益事件。LLM方法生成包含情绪极性、看涨度、讽刺可能性和主题相关性的多维情感表示,而基线方法采用VADER词典模型。通过领先滞后相关性分析、OLS回归、ROC-AUC方向分类及分位数早期预警框架评估两种方法。结果表明,LLM衍生指标提供更丰富的多维表征,且具有更强的资产特定统计结构;但其与市场走势的关系在不同资产间存在异质性,说明更高的语言表达能力并不必然带来零售驱动波动环境下的稳定预测性能。
原文摘要 · Abstract (English)
This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。