arXiv:2412.02605cs.CLcs.LG2024-12

用稀疏自编码器挖掘公司描述中的可解释特征,提升投资相似性判断。

Interpretable Company Similarity with Sparse Autoencoders

  • 通过稀疏自编码器将公司描述分解为可解释的语义特征。
  • 在月度收益相关性和协整交易策略中,表现优于行业分类和嵌入向量。
  • 特征可解释性强,适合金融风控与组合管理等高要求场景。

确定公司相似性是金融领域的重要任务,支撑风险管控、对冲和投资组合多元化。从业者常依赖SIC和GICS等行业分类来判断相似性,前者由美国证券交易委员会(SEC)使用,后者被投资界广泛采用。但这些分类粒度不足且需定期更新。因此,有研究提出使用公司描述的嵌入聚类作为替代方案,但词元嵌入缺乏可解释性,限制了其在高风险场景中的应用。稀疏自编码器(SAEs)在提升大语言模型(LLM)内部表示可解释性方面表现优异,能将模型激活分解为可解释特征,并捕捉公司描述的深层表征,而不仅是语义相似性。本文将SAEs应用于公司描述,生成具有实际意义的股票聚类。我们以SIC码、行业代码和嵌入向量为基准进行对比,结果表明,SAE特征在捕捉公司基本面特征方面优于传统分类与嵌入。这一优势体现在更高的月度收益率相关性(作为相似性代理指标)以及在协整交易策略中产生更高夏普比率,表明公司间存在更深层次的基本面相似性。最后,我们验证了聚类的可解释性,证明稀疏特征能形成简洁直观的解释。

原文摘要 · Abstract (English)

Determining company similarity is a vital task in finance, underpinning risk management, hedging, and portfolio diversification. Practitioners often rely on sector and industry classifications such as SIC and GICS codes to gauge similarity, the former being used by the U.S. Securities and Exchange Commission (SEC), and the latter widely used by the investment community. Since these classifications lack granularity and need regular updating, using clusters of embeddings of company descriptions has been proposed as a potential alternative, but the lack of interpretability in token embeddings poses a significant barrier to adoption in high-stakes contexts. Sparse Autoencoders (SAEs) have shown promise in enhancing the interpretability of Large Language Models (LLMs) by decomposing Large Language Model (LLM) activations into interpretable features. Moreover, SAEs capture an LLM's internal representation of a company description, as opposed to semantic similarity alone, as is the case with embeddings. We apply SAEs to company descriptions, and obtain meaningful clusters of equities. We benchmark SAE features against SIC-codes, Industry codes, and Embeddings. Our results demonstrate that SAE features surpass sector classifications and embeddings in capturing fundamental company characteristics. This is evidenced by their superior performance in correlating logged monthly returns - a proxy for similarity - and generating higher Sharpe ratios in co-integration trading strategies, which underscores deeper fundamental similarities among companies. Finally, we verify the interpretability of our clusters, and demonstrate that sparse features form simple and interpretable explanations for our clusters.

公司相似性稀疏自编码器金融分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。