arXiv:2502.15072stat.MLcs.LG2025-02

改进树模型分叉策略,更精准识别高风险人群

Policy-Oriented Binary Classification: Improving (KD-)CART Final Splits for Subpopulation Targeting

  • 提出MDFS分叉方法,基于概率距离最大化优化分层规则
  • 在模拟与真实数据中均显著提升对脆弱群体的识别率
  • 适合政策制定者用于精准靶向高风险子人群

政策制定者常利用递归二元分裂规则,根据二元结果对人群进行划分,并针对事件发生概率超过阈值的子群体实施干预。这类问题称为潜在概率分类(LPC)。实践中普遍采用分类与回归树(CART)方法。我们证明,在LPC背景下,经典CART及知识蒸馏方法(学生模型为CART,简称KD-CART)存在次优性。本文提出最大距离最终分叉(MDFS),在唯一交集假设下,其生成的分叉规则严格优于CART/KD-CART。MDFS能识别出唯一最优分叉规则,具有一致性,且比CART/KD-CART更有效定位脆弱子群体。为放宽唯一交集假设,进一步提出惩罚最终分叉(PFS)和加权经验风险最终分叉(wEFS)。大量模拟研究显示,所提方法总体上优于CART/KD-CART。在真实数据应用中,MDFS生成的政策也更有效地聚焦于更脆弱的人群。

原文摘要 · Abstract (English)

Policymakers often use recursive binary split rules to partition populations based on binary outcomes and target subpopulations whose probability of the binary event exceeds a threshold. We call such problems Latent Probability Classification (LPC). Practitioners typically employ Classification and Regression Trees (CART) for LPC. We prove that in the context of LPC, classic CART and the knowledge distillation method, whose student model is a CART (referred to as KD-CART), are suboptimal. We propose Maximizing Distance Final Split (MDFS), which generates split rules that strictly dominate CART/KD-CART under the unique intersect assumption. MDFS identifies the unique best split rule, is consistent, and targets more vulnerable subpopulations than CART/KD-CART. To relax the unique intersect assumption, we additionally propose Penalized Final Split (PFS) and weighted Empirical risk Final Split (wEFS). Through extensive simulation studies, we demonstrate that the proposed methods predominantly outperform CART/KD-CART. When applied to real-world datasets, MDFS generates policies that target more vulnerable subpopulations than the CART/KD-CART.

决策树政策优化子群体识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。