SAGA用自适应生成架构预测长期收入,精度远超传统模型。
SAGA: A Sequence-Adaptive Generative Architecture for Multi-Horizon Probabilistic Forecasting with Adaptive Temporal Conformal Prediction
- 基于解码器仅的Transformer处理不规则面板数据,支持多时域概率预测。
- 十年预测CRPS降低31.9%,二十年预测MAE降低37.7%,优于经典参数模型。
- 结合分拆共形校准,个体预测区间覆盖率接近名义水平,适合政策评估场景。
财政与央行使用的微观模拟模型依赖参数化收入过程,仅捕捉条件分布的一阶和二阶矩,忽略长期非线性结构。本文提出SAGA——一种用于不规则表格型面板序列的解码器仅变压器,搭配分拆共形校准封装,实现具有有限样本边际覆盖保证的个体级预测区间。模型在1990至2022年瑞典LISA登记数据上训练,包含2,143,817名个体与61,284,903人年数据,可预测未来1至30年的年度劳动收入,并通过蒙特卡洛聚合得到现值折现终身收入分布。相较于经典的Guvenen、Karahan、Ozkan、Song参数模型及表格式与递归基线,SAGA在十年预测中连续排名概率评分降低31.9%,二十年预测平均绝对误差降低37.7%。共形区间在整体上达到名义覆盖率的0.4个百分点以内,最差人口子群内为2.4个百分点以内。重构的终身收入基尼系数为0.327,接近部分观测真实值0.341,优于GKOS估计值0.378。模型权重、校准表与合成等价数据集已公开,供在受保护的SCB MONA环境外复现。
原文摘要 · Abstract (English)
Microsimulation models used by ministries of finance and central banks rely on parametric processes for lifetime earnings that capture only first and second moments of the conditional distribution and miss long-range nonlinear structure. We propose SAGA, a decoder-only transformer for irregular tabular panel sequences, paired with a split conformal calibration wrapper that delivers individual-level prediction intervals with finite-sample marginal coverage guarantees. Trained on the longitudinal Swedish LISA register over 1990 to 2022, comprising 2,143,817 individuals and 61,284,903 person-years, the model forecasts annual labor earnings at horizons of one to thirty years and aggregates them by Monte Carlo into present-discounted lifetime earnings distributions. Against the canonical Guvenen, Karahan, Ozkan, and Song parametric process and tabular and recurrent baselines, SAGA reduces continuous ranked probability score by 31.9 percent at the ten-year horizon and mean absolute error by 37.7 percent at the twenty-year horizon. Conformal intervals achieve nominal coverage to within 0.4 percentage points marginally and within 2.4 percentage points on the worst-case demographic subgroup. The reconstructed lifetime earnings Gini coefficient is 0.327 against the partially observed truth of 0.341 and the GKOS estimate of 0.378. Model weights, calibration tables, and a synthetic equivalent dataset are released for replication outside the protected SCB MONA environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。