用语义匹配和多层级分类构建金融文本-时序数据集,提升股价预测效果。
FinTexTS: Financial Text-Paired Time-Series Dataset via Semantic-Based and Multi-Level Pairing
- 基于公司财报上下文与嵌入匹配,精准关联新闻与股价数据。
- 按宏观、行业、相关公司、目标公司四级分类,实现多层级新闻配对。
- 适用于量化研究、金融预测模型开发人员,尤其关注文本-时序融合者。
金融领域存在多种重要时序问题。近年来,结合文本与数值信息的时序分析方法受到广泛关注。为此,已有大量工作致力于构建金融领域的文本-时序数据集。然而,金融市场具有复杂的相互依赖性:公司股价不仅受自身事件影响,还受其他公司及宏观经济因素影响。现有基于关键词匹配的文本-时序配对方法难以捕捉此类复杂关系。为解决此问题,本文提出一种基于语义的多层级配对框架。首先,从美国证监会(SEC)文件中提取目标公司的特定上下文,并使用嵌入式匹配机制检索语义相关的新闻文章。其次,利用大语言模型(LLMs)将新闻文章分为四个层级(宏观级、行业级、相关公司级、目标公司级),实现多层级新闻与目标公司的配对。基于公开新闻数据集,我们构建了FinTexTS——一个大规模文本-时序股价数据集。在FinTexTS上的实验表明,该语义驱动的多层级配对策略在股价预测任务中表现优异。此外,将该方法应用于专有且精心筛选的新闻源,可生成更高品质的配对数据,并进一步提升股价预测性能。
原文摘要 · Abstract (English)
The financial domain involves a variety of important time-series problems. Recently, time-series analysis methods that jointly leverage textual and numerical information have gained increasing attention. Accordingly, numerous efforts have been made to construct text-paired time-series datasets in the financial domain. However, financial markets are characterized by complex interdependencies, in which a company's stock price is influenced not only by company-specific events but also by events in other companies and broader macroeconomic factors. Existing approaches that pair text with financial time-series data based on simple keyword matching often fail to capture such complex relationships. To address this limitation, we propose a semantic-based and multi-level pairing framework. Specifically, we extract company-specific context for the target company from SEC filings and apply an embedding-based matching mechanism to retrieve semantically relevant news articles based on this context. Furthermore, we classify news articles into four levels (macro-level, sector-level, related company-level, and target company-level) using large language models (LLMs), enabling multi-level pairing of news articles with the target company. Applying this framework to publicly-available news datasets, we construct FinTexTS, a new large-scale text-paired stock price dataset. Experimental results on FinTexTS demonstrate the effectiveness of our semantic-based and multi-level pairing strategy in stock price forecasting. In addition to publicly-available news underlying FinTexTS, we show that applying our method to proprietary yet carefully curated news sources leads to higher-quality paired data and improved stock price forecasting performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。