用上下文感知NLP从投资会议中提取预测金融信号的特征
Converting Expert Deliberation into Financial Signals Through A Context-Aware NLP Pipeline

- 通过大模型分段并标注资产类别,构建情感极性和提及频率等结构化特征
- 组合嵌入与工程特征的模型准确率达73%,优于始终选择股票的60.4%基准
- 专家讨论中的情感信息比提及频次更具预测价值,适合量化研究者参考
我们提出CDSP(上下文条件的审议信号管道),将投资委员会会议记录转化为结构化预测特征。该方法将会议文本分段,利用大语言模型(LLM)分配资产类别标签,将金融关键词映射到预定义标签体系,并构建情感极性与提及频率等互补特征。该特征工程框架应用于涵盖48次月度会议的数据集,用于预测全球股市下月表现是否优于全球债券。实验显示,使用工程特征、原始文本、句子嵌入及组合表示的模型准确率在62%至73%之间,而始终选择股票的基准准确率为60.4%。最佳模型(73%准确率)结合句子嵌入与工程特征,达到0.73 F1得分(未达统计显著)。多个分类体系中,情感信号强于提及频率。结果表明,专家讨论可能包含前瞻性信息,可通过上下文感知NLP有效提取。
原文摘要 · Abstract (English)
We introduce the CDSP (context-conditional deliberation signal pipeline), converting an investment committee's meeting transcripts into structured predictive features. CDSP segments the meeting transcripts into topical chunks, assigns asset-class context labels using a large language model (LLM), maps financial keywords to a pre-determined taxonomy of labels, and constructs complementary features: sentiment polarity and mention frequency. This feature engineering framework is applied to a dataset spanning 48 monthly committee meetings to predict if global equities will perform better or worse than global bonds in the following month. In experiments with engineered features, raw transcript text, sentence embeddings, and combined representations, the prediction accuracy ranges from 62% to 73%, compared to always choosing stocks, which outperforms bonds 60.4% of the time. The best (73% accurate) model combines sentence embeddings with engineered CDSP features, achieving a 0.73 F1 score (although this is not statistically significant compared to always choosing stocks). Sentiment carries a stronger signal than mention frequency for several taxonomy categories. These findings suggest that experts' deliberations may contain forward-looking information that context-aware NLP can extract.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。