arXiv:2609.03106cs.LGq-fin.RM2026-09被引 1

研究保险数据中模型性能随数据量增长的规律,发现特定架构更优。

Scaling Laws, Tabular Data and Actuarial Ratemaking Models

论文配图:Scaling Laws, Tabular Data and Actuarial Ratemaking Models
图 1 · 摘自论文原文
  • 在真实车险数据上测试不同模型,观察其性能随数据量变化趋势。
  • TabM 模型的数据扩展能力远超传统 Transformer 和 MLP,提升显著。
  • 适合关注保险建模、数据规模与模型选择的从业者参考。

现代深度学习中的缩放定律描述了模型容量、训练数据和计算资源增加时,保留损失如何改善,常呈现幂律趋势。我们探究此类缩放规律是否适用于保险精算中的列式数据,这类数据具有异质性和噪声,且经典模型如广义线性模型(GLMs)仍是强基线。基于一个真实的车险保单组合,我们在不同数据比例下训练多种模型家族,并使用多个随机种子进行评估,采用样本外泊松偏差(Poisson deviance)作为损失指标,该指标越低表示预测拟合越好。结果表明,所有模型家族在数据增加时性能均提升,但缩放指数差异明显:TabM 在数据扩展方面表现显著优于纯监督的表格型 Transformer 和标准 MLP 基线。仅通过扩大参数量的 Transformer 变体表现出弱缩放能力,除非引入额外归纳偏置(如 TabM 式适配或自监督)。这些结果为不同数据规模下的模型选择提供了量化依据,表明有效缩放依赖于架构设计与损失函数目标,单纯增大 Transformer 规模带来的收益有限。

原文摘要 · Abstract (English)

Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends. We investigate whether analogous scaling regularities arise in actuarial ratemaking, where data are tabular, heterogeneous, and noisy, and where classical models such as GLMs remain strong baselines. Using a real-world motor insurance portfolio, we train models from different families across increasing fractions of the training data and multiple random seeds, evaluating out-of-sample Poisson deviance, a likelihood-based loss for Poisson count predictions in which lower values indicate better held-out fit. We find that all model families improve with additional data, but scaling exponents differ substantially: TabM exhibits markedly stronger data scaling than purely supervised tabular Transformers and standard MLP baselines. Transformer variants show weak parameter scaling unless augmented with additional inductive biases (TabM-style adaptation or self-supervision). These results provide quantitative guidance on model selection by data regime and suggest that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.

精算建模表格数据缩放定律保险科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。