用注意力模型改进股票定价,能更好捕捉时间依赖与市场风险。
Is attention truly all we need? An empirical study of asset pricing in pretrained RNN sparse and global attention models
- 对比多种注意力机制,结合因果掩码避免未来信息泄露
- 新冠期间全球自注意模型年化索提诺比率达2.0,滑动窗口稀疏模型为1.80
- 适合关注金融时序建模与风险对冲的量化研究者
本研究首次将主流注意力机制(如加性注意力、Luong三类注意力、全局自注意力、滑动窗口稀疏注意力)应用于前420只大型美股的实证资产定价。相比传统机器学习方法,这些预训练RNN注意力模型克服了对时间依赖性捕捉不足和短期记忆缺陷,并通过强制因果掩码解决了经典Transformer中被忽视的未来数据泄露问题。模型还考虑了资产定价数据的时间稀疏特性,通过简化结构缓解过拟合。在疫情前、疫情期间及疫情后一年三个阶段测试稳定性。结果显示,在价值加权投资组合回测中,全局自注意力与滑动窗口稀疏注意力模型在新冠期间分别实现2.0和1.80的年化索提诺比率(静态交易成本下),且后者在不同市值股票上表现更稳定。
原文摘要 · Abstract (English)
This study investigates the pre-trained RNN attention models with the mainstream attention mechanisms, such as additive attention, Luong's three attentions, global self-attention and sliding window sparse attention, for the empirical asset pricing research on the top 420 large-cap US stocks. This is the first paper on the large-scale state-of-the-art (SOTA) attention mechanisms applied in the asset pricing context. They overcome the limitations of the traditional machine learning-based asset pricing, such as mis-capturing the temporal dependency and short memory. Moreover, the enforced causal masks in the attention mechanisms address the future data leaking issue ignored by the more advanced attention-based models, such as the classic Transformer. The proposed attention models also consider the temporal sparsity characteristic of asset pricing data and mitigate potential overfitting issues by deploying the simplified model structures. This provides some insights for future empirical economic research. All models are examined in three periods, which cover pre-COVID-19, COVID-19 and one year post-COVID-19, for testing the stability of these models under extreme market conditions. The study finds that in value-weighted portfolio back testing, the global self-attention model and the sliding window sparse attention model exhibit excellent capabilities in deriving the absolute returns and hedging downside risks, while they achieve an annualized Sortino ratio of 2.0 and 1.80 respectively in the period with COVID-19 in the static transaction cost scenario. Moreover, the sliding window sparse attention model performs more stably than the global self-attention model from the perspective of absolute portfolio returns with respect to the size of stocks' market capitalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。