arXiv:2410.04684stat.APcs.LG2024-10被引 1

融合保险理赔文本与金额,提升风险分类与预测精度

Combining Structural and Unstructured Data: A Topic-based Finite Mixture Model for Insurance Claim Prediction

  • 用主题模型关联理赔文本与损失金额,构建联合概率框架
  • 在真实数据上实现更精准的理赔聚类与金额预测
  • 适合需要可解释性理赔分析的保险公司和精算师

保险理赔金额建模与风险等级分类是关键但具有挑战性的任务。传统预测模型常忽略理赔描述中的有价值信息。本文提出一种新方法,构建一个联合混合模型,同时整合理赔描述与理赔金额。该方法建立文本描述与损失金额之间的概率关联,提升了理赔聚类与预测的准确性。在所提模型中,潜在主题/分量指标同时作为理赔描述的主题内容与损失分布的组成部分的代理变量。具体而言,在给定主题/分量指标的条件下,理赔描述服从多项式分布,而理赔金额服从对应分量的损失分布。我们提出了两种模型校准方法:用于最大后验估计的EM算法,以及用于后验分布推断的MH-within-Gibbs采样算法。实证研究表明,所提方法有效,能够提供可解释的理赔聚类与预测结果。

原文摘要 · Abstract (English)

Modeling insurance claim amounts and classifying claims into different risk levels are critical yet challenging tasks. Traditional predictive models for insurance claims often overlook the valuable information embedded in claim descriptions. This paper introduces a novel approach by developing a joint mixture model that integrates both claim descriptions and claim amounts. Our method establishes a probabilistic link between textual descriptions and loss amounts, enhancing the accuracy of claims clustering and prediction. In our proposed model, the latent topic/component indicator serves as a proxy for both the thematic content of the claim description and the component of loss distributions. Specifically, conditioned on the topic/component indicator, the claim description follows a multinomial distribution, while the claim amount follows a component loss distribution. We propose two methods for model calibration: an EM algorithm for maximum a posteriori estimates, and an MH-within-Gibbs sampler algorithm for the posterior distribution. The empirical study demonstrates that the proposed methods work effectively, providing interpretable claims clustering and prediction.

理赔预测主题模型联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。