arXiv:2605.22892q-fin.RMcs.LG2026-05

TabPFN在车险定价中表现不如传统模型,推理慢且对数据量敏感。

Is TabPFN the Silver Bullet for Insurance Pricing?

  • 用预训练的TabPFN通过上下文学习直接推理,无需调参
  • 在两个公开数据集上表现不及GLM和XGBoost,平均误差更高
  • 适合小样本场景,但当前不适用于主流保险定价

非寿险定价中,索赔频率与严重性的建模通常依赖广义线性模型(GLM),梯度提升机是主要的机器学习替代方案。表格式基础模型(TFMs)提出了一种根本不同的推理范式:通过在大量合成数据集上预训练,实现无需特定数据集拟合或超参数调优的上下文学习推理。本文首次对TabPFN在车险定价中的表现进行实证评估,将其与GLM和XGBoost在两个公开的MTPL数据集上进行对比。结果表明,TabPFN并未持续优于现有基线模型,在推理时间上显著更长,且对上下文训练集大小敏感。尽管表格式基础模型在数据稀缺场景下前景可期,但其当前性能尚不足以替代成熟的精算方法。

原文摘要 · Abstract (English)

Modelling claim frequency and severity for non-life insurance pricing predominantly relies on generalised linear models, with gradient-boosted machines as the leading machine learning alternative. Tabular foundation models (TFMs) present a fundamentally different inference paradigm. By pre-training on large collections of synthetic datasets, TFMs enable inference on new data through in-context learning, without any dataset-specific fitting or hyperparameter tuning. This paper presents a first empirical evaluation of TabPFN for motor insurance pricing, benchmarking it against GLM and XGBoost on two publicly available MTPL datasets. Our results show that TabPFN does not consistently outperform established baselines, exhibits substantially longer inference times, and is sensitive to the size of the in-context training set. While tabular foundation models represent a promising direction, particularly in data-scarce settings, their current performance does not offer a viable replacement for established actuarial methods.

保险定价表格式模型上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。