arXiv:2511.15446stat.MLcs.LG2025-11被引 3

改进了吉尼分数在存在风险排名并列和案例权重时的使用方法。

Gini Score under Ties and Case Weights

  • 提出并修正了吉尼分数在排名并列情况下的计算方式。
  • 引入案例权重后,仍保持吉尼分数对风险排序的有效性。
  • 适用于保险精算中存在数据权重和重复评分的实际场景。

吉尼分数是统计建模和机器学习中用于模型验证与选择的常用工具,是一种基于排序的评分方法,可用于评估风险排序。该分数在二分类场景中广泛应用,且与受试者工作特征(ROC)或曲线下面积(AUC)等价。在精算文献中,这一基于排序的分数被扩展至一般实值随机变量,采用洛伦兹曲线与集中曲线。尽管这些早期概念假设风险排序由连续分布生成,本文探讨了在风险排序存在并列时如何应用吉尼分数,并进一步将吉尼分数适应于精算中常见的案例权重情形。

原文摘要 · Abstract (English)

The Gini score is a popular tool in statistical modeling and machine learning for model validation and model selection. It is a purely rank based score that allows one to assess risk rankings. The Gini score for statistical modeling has mainly been used in a binary context, in which it has many equivalent reformulations such as the receiver operating characteristic (ROC) or the area under the curve (AUC). In the actuarial literature, this rank based score for binary responses has been extended to general real-valued random variables using Lorenz curves and concentration curves. While these initial concepts assume that the risk ranking is generated by a continuous distribution function, we discuss in this paper how the Gini score can be used in the case of ties in the risk ranking. Moreover, we adapt the Gini score to the common actuarial situation of having case weights.

吉尼分数风险排序精算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。