用对比学习挖掘财经新闻与市场走势的语义关联,提升预测准确率7%。
Contrastive Similarity Learning for Market Forecasting: The ContraSim Framework
- 通过加权新闻增强和加权自监督对比学习构建语义嵌入空间
- 在华尔街日报数据上使分类准确率提升7%,且不依赖标签
- 可找出与当前新闻相似的历史行情日,辅助分析师判断趋势
我们提出对比相似性空间嵌入算法(ContraSim),用于揭示每日财经新闻与市场走势之间的全局语义关系。该框架包含两个关键阶段:(I) 加权新闻增强,生成增强后的财经新闻并附带细粒度语义相似度评分;(II) 加权自监督对比学习(WSSCL),基于相似度度量构建优化的加权嵌入空间,使语义相近的新闻聚类。实证结果表明,将ContraSim特征融入金融预测任务后,基于华尔街日报(WSJ) headlines的分类准确率提升7%。进一步的信息密度分析显示,ContraSim构建的相似性空间能自然聚类出市场走势方向一致的日子,说明其捕捉了无需标签的市场动态。此外,该方法可识别出与当前新闻高度相似的历史新闻日,为分析师提供参考过去事件预测未来趋势的可操作洞见。
原文摘要 · Abstract (English)
We introduce the Contrastive Similarity Space Embedding Algorithm (ContraSim), a novel framework for uncovering the global semantic relationships between daily financial headlines and market movements. ContraSim operates in two key stages: (I) Weighted Headline Augmentation, which generates augmented financial headlines along with a semantic fine-grained similarity score, and (II) Weighted Self-Supervised Contrastive Learning (WSSCL), an extended version of classical self-supervised contrastive learning that uses the similarity metric to create a refined weighted embedding space. This embedding space clusters semantically similar headlines together, facilitating deeper market insights. Empirical results demonstrate that integrating ContraSim features into financial forecasting tasks improves classification accuracy from WSJ headlines by 7%. Moreover, leveraging an information density analysis, we find that the similarity spaces constructed by ContraSim intrinsically cluster days with homogeneous market movement directions, indicating that ContraSim captures market dynamics independent of ground truth labels. Additionally, ContraSim enables the identification of historical news days that closely resemble the headlines of the current day, providing analysts with actionable insights to predict market trends by referencing analogous past events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。