arXiv:2505.14420q-fin.CPcs.CL2025-05ACL被引 3

用稀疏自编码器从财报文本中提取关键特征,提升盈利预测准确率。

SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection

  • 通过稀疏自编码器将语言模型的密集表示分解为可解释的稀疏特征。
  • 在三个金融数据集上,预测性能显著优于基线方法。
  • 适合关注金融文本分析与模型可解释性的研究者使用。

从财务文档(如业绩电话会议、监管文件和财经新闻)中预测盈利意外,在金融经济学中日益重要。然而,这些文档通常超过5000字,存在大量冗余信息和行业术语,给语言模型带来挑战。本文提出SAE-FiRE(稀疏自编码器用于金融表征增强)框架,通过提取关键信息并消除冗余来应对这一问题。SAE-FiRE利用稀疏自编码器(SAEs)将大型语言模型的密集神经表示分解为可解释的稀疏成分,并结合方差分析F检验和树模型重要性评分等统计方法,筛选出前k个最具判别力的维度用于分类。通过系统性过滤可能引发过拟合的噪声,实现更鲁棒且泛化能力更强的预测。在三个金融数据集上的实验表明,SAE-FiRE显著优于基线方法。

原文摘要 · Abstract (English)

Predicting earnings surprises from financial documents, such as earnings conference calls, regulatory filings, and financial news, has become increasingly important in financial economics. However, these financial documents present significant analytical challenges, typically containing over 5,000 words with substantial redundancy and industry-specific terminology that creates obstacles for language models. In this work, we propose the SAE-FiRE (Sparse Autoencoder for Financial Representation Enhancement) framework to address these limitations by extracting key information while eliminating redundancy. SAE-FiRE employs Sparse Autoencoders (SAEs) to decompose dense neural representations from large language models into interpretable sparse components, then applies statistical feature selection methods, including ANOVA F-tests and tree-based importance scoring, to identify the top-k most discriminative dimensions for classification. By systematically filtering out noise that might otherwise lead to overfitting, we enable more robust and generalizable predictions. Experimental results across three financial datasets demonstrate that SAE-FiRE significantly outperforms baseline approaches.

金融预测自编码器特征选择语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。