arXiv:2607.14174cs.LGq-fin.CP2026-07

研究10-K文件文本与风险因子情感的预测价值,发现不同层级下文本范围影响效果

How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment

论文配图:How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment
图 1 · 摘自论文原文
  • 用监督学习方法在10-K全文和风险因子部分训练情感分数
  • 全文件文本在行业和投资组合层级更准,个体公司层级反而风险因子部分更优
  • 揭示了文档量与训练信号量的交互影响,适合金融风控与量化研究者

财务情感提取长期依赖新闻文本和基于收益标签的监督方法,对10-K文件——尤其是其第1A项风险因子部分——关注不足。本文将监督词典学习方法拓展至10-K及其第1A项内容,分别以收益和波动率为标签,在行业、投资组合和单家公司三个聚合层级上训练情感得分。基于94家纳斯达克100科技企业2006—2023年的1,383份文件,评估了12种情感指标在分类准确率、与实际市场结果相关性及词法内容上的表现。结果显示,全文件文本在行业和投资组合层级对两类目标均更精准,但在个体公司层级反被第1A项内容超越;这一现象归因于文档体量与可用独立训练信号在不同聚合层级的交互作用。基于Loughran-McDonald词典的基线模型在所有层级均与股价呈显著负相关,凸显监督方法在监管披露文本中的优势。这些发现及其设计启示,为后续更大规模多源系统的情感生成方法奠定了基础。

原文摘要 · Abstract (English)

Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored. We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation: sector, portfolio, and individual firm. Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006--2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, correlation with realised market outcomes, and qualitative lexical content. Full-filing text produces more accurate sentiment at the sector and portfolio level for both targets, but this reverses at the individual-firm level, where the narrower Item 1A section performs better -- an effect we attribute to the interaction between document volume and the amount of independent training signal available at each level of aggregation. A Loughran-McDonald dictionary baseline is consistently, strongly negatively correlated with price at every level tested, underscoring the value of a supervised approach for regulatory disclosure text. These findings, and the design choices they motivate, establish the sentiment-generation methodology underlying a subsequent, larger-scale, multi-source system.

金融情感10-K分析风险因子监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。