用机制模型+数据模型结合,更准预测公司增长
Predicting Company Growth using Scaling Theory informed Machine Learning
- 融合尺度理论模型与机器学习,分离趋势与波动
- 在3万多家公司数据上,预测准确率显著提升
- 适合金融预测、企业分析领域研究者参考
预测公司增长是关键但极具挑战的任务,因观测到的动态同时包含结构性增长趋势与剧烈波动。本文提出一种基于尺度理论的机器学习框架(STIML),将基于尺度的增长模型与数据驱动的预测模型结合,分别捕捉机制驱动的平均趋势和残差波动。基于1950至2019年北美31,553家公司的Compustat年度财务数据,我们将增长模型拓展至多个财务指标,并在16个目标变量上评估STIML,相较于仅用增长模型或纯数据驱动的基线模型表现更优。结果表明,公司增长中趋势驱动与波动驱动的可预测性存在明显分离,其相对重要性强烈依赖于公司规模与波动水平。可解释性分析显示,STIML能捕捉多变量间非简单自相关依赖关系,而宏观经济变量对预测性能的贡献平均较低。此外,发现尺度模型忽略的非对称偏离反而蕴含结构化且可学习信号,为改进机制模型提供了方向。
原文摘要 · Abstract (English)
Predicting company growth is a critical yet challenging task because observed dynamics blend an underlying structural growth trend with volatile fluctuations. Here, we propose a Scaling-Theory-Informed Machine Learning (STIML) framework that integrates a scaling-based growth model to capture the mechanism-driven average trend, together with a data-driven forecasting model to learn the residual fluctuations. Using Compustat annual financial statement data (1950--2019) for 31,553 North American companies, we extend the growth model beyond assets to multiple financial indicators, and evaluate STIML against growth model-only and purely data-driven baselines. Across 16 target variables, we show that company growth exhibits a clear separation between trend-driven predictability and fluctuation-driven predictability, with their relative importance depending strongly on company size and volatility. Interpretability analyses further show that STIML captures multivariate dependencies beyond simple autocorrelation, and that macroeconomic variables contribute significantly less to predictive performance on average. Moreover, we find the scaling-based growth model overlooks asymmetric deviations, which instead contain the structured and learnable signals, suggesting a path to refine mechanistic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。