提出可灵活遗忘的树模型训练方法,提升隐私保护效率。
FUTURE: Flexible Unlearning for Tree Ensemble
- 将遗忘建模为梯度优化问题,用概率近似解决树结构不可导难题
- 在真实数据集上实现高效且显著的遗忘效果,性能优于现有方法
- 适合需要隐私保护的医疗、金融等高敏感领域应用
树集成模型在分类任务中表现优异,广泛应用于生物信息学、金融和医疗诊断等领域。随着数据隐私和‘被遗忘权’受重视,已有多种遗忘算法用于让树集成模型删除敏感信息。但现有方法通常针对特定模型或依赖离散树结构,难以推广到复杂集成模型,且在大规模数据上效率低下。为此,我们提出 FUTURE,一种新型树集成遗忘算法。将样本遗忘问题建模为基于梯度的优化任务,通过在优化框架中采用概率模型近似,克服树集成不可导的难点,实现端到端高效遗忘。在多个真实数据集上的实验表明,FUTURE 能取得显著且成功的遗忘效果。
原文摘要 · Abstract (English)
Tree ensembles are widely recognized for their effectiveness in classification tasks, achieving state-of-the-art performance across diverse domains, including bioinformatics, finance, and medical diagnosis. With increasing emphasis on data privacy and the \textit{right to be forgotten}, several unlearning algorithms have been proposed to enable tree ensembles to forget sensitive information. However, existing methods are often tailored to a particular model or rely on the discrete tree structure, making them difficult to generalize to complex ensembles and inefficient for large-scale datasets. To address these limitations, we propose FUTURE, a novel unlearning algorithm for tree ensembles. Specifically, we formulate the problem of forgetting samples as a gradient-based optimization task. In order to accommodate non-differentiability of tree ensembles, we adopt the probabilistic model approximations within the optimization framework. This enables end-to-end unlearning in an effective and efficient manner. Extensive experiments on real-world datasets show that FUTURE yields significant and successful unlearning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。