arXiv:2509.02171stat.MLcs.LG2025-09

用多重插补法生成保险定价用合成数据,效果优于深度生成模型。

Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders

  • 采用多重插补法(MICE)生成保险数据,无需复杂训练。
  • 合成数据保留原始变量分布与协变量关系,表现接近真实数据。
  • 适合预算有限或需快速部署的保险建模团队使用。

保险精算依赖高质量数据,但获取新数据成本高且存在隐私问题。本文探索合成数据生成作为解决方案,对比了传统的生成模型(如变分自编码器、条件表格式生成对抗网络)与基于多重插补法(MICE)的新方法。在开源数据集上的比较研究表明,MICE方法能有效保留原始变量的边际分布及协变量间的多变量关系,且基于合成数据训练的广义线性模型(GLMs)与真实数据训练结果高度一致。同时,该方法实现复杂度远低于深度生成模型,且对原始数据进行合成数据增强后,可提升预测索赔次数的GLM性能。结果表明,MICE方法在生成高质量表格数据方面具有显著潜力。

原文摘要 · Abstract (English)

Actuarial ratemaking depends on high-quality data, yet access to such data is often limited by the cost of obtaining new data, privacy concerns, etc. In this paper, we explore synthetic-data generation as a potential solution to these issues. In addition to generative methods previously studied in the actuarial literature, we explore and benchmark another class of approaches based on Multivariate Imputation by Chained Equations (MICE). In a comparative study using an open-source dataset, MICE-based models are evaluated against other generative models like Variational Autoencoders and Conditional Tabular Generative Adversarial Networks. We assess how well synthetic data preserves the original marginal distributions of variables as well as the multivariate relationships among covariates. The consistency between Generalized Linear Models (GLMs) trained on synthetic data with GLMs trained on the original data is also investigated. Furthermore, we assess the ease of use of each generative approach and study the impact of generically augmenting original data with synthetic data on the performance of GLMs for predicting claim counts. Our results highlight the potential of MICE-based methods in creating high-fidelity tabular data while offering lower implementation complexity compared to deep generative models.

合成数据保险精算多重插补生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。