arXiv:2601.07637q-fin.RMcs.LG2026-01被引 2

用强化学习动态更新理赔预估,更准且能用未结案数据。

Reinforcement Learning for Micro-Level Claims Reserving

  • 将每笔理赔视为马尔可夫决策过程,逐步优化预估金额。
  • 在未结案样本上训练,提升对早期理赔的预测准确率。
  • 适合精算师用于处理复杂、长期的保险理赔建模。

未决理赔负债会随理赔发展不断调整,但多数现代精算模型仅基于已结算案件进行一次性预测。本文将个体理赔预测建模为理赔级马尔可夫决策过程,代理通过连续动作和奖励机制,在理赔发展过程中逐步修正未决负债(OCL)估计值。该强化学习方法可利用所有观测到的理赔轨迹,包括仍在进展中的案件,避免仅依赖最终结果带来的样本量减少与选择偏差。研究还引入了新案初始化、滚动结算调参及重要性加权机制,以缓解大额理赔罕见导致的组合层面低估问题。在CAS与SPLICE合成通用保险数据集上,所提出的软演员-评论家算法在理赔级精度与整体OCL表现上均具竞争力,尤其在驱动主要负债的早期未成熟理赔中表现突出。

原文摘要 · Abstract (English)

Outstanding claim liabilities are revised repeatedly as claims develop, yet most modern reserving models are trained as one-shot predictors and typically learn only from settled claims. We formulate individual claims reserving as a claim-level Markov decision process in which an agent sequentially updates outstanding claim liability (OCL) estimates over development, using continuous actions and a reward design that balances accuracy with stable reserve revisions. A key advantage of this reinforcement learning (RL) approach is that it can learn from all observed claim trajectories, including claims that remain open at valuation, thereby avoiding the reduced sample size and selection effects inherent in supervised methods trained on ultimate outcomes only. We also introduce practical components needed for actuarial use -- initialisation of new claims, temporally consistent tuning via a rolling-settlement scheme, and an importance-weighting mechanism to mitigate portfolio-level underestimation driven by the rarity of large claims. On CAS and SPLICE synthetic general insurance datasets, the proposed Soft Actor-Critic implementation delivers competitive claim-level accuracy and strong aggregate OCL performance, particularly for the immature claim segments that drive most of the liability.

强化学习理赔预测精算建模动态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。