arXiv:2501.11823cs.LGcs.AI2025-01被引 5

提出新方法让大图模型高效删掉指定数据,兼顾遗忘效果与推理能力。

Toward Scalable Graph Unlearning: A Node Influence Maximization based Approach

  • 用节点影响力最大化思路解耦图中节点影响,提升可扩展性。
  • 在14个数据集上验证,包括ogbn-papers100M,显著增强遗忘能力。
  • 可插拔集成到多数图学习方法中,适合大规模图隐私保护场景。

机器遗忘作为提升模型鲁棒性与数据隐私的关键技术,在当前流行的图挖掘应用中备受关注,尤其在大规模图场景下。然而,现有图遗忘(GU)方法因网络规模下图元素间复杂交互面临两大挑战:(1) 梯度驱动的节点纠缠阻碍对遗忘请求的完全知识清除;(2) 网络级图元素(超十亿量级)带来不可避免的可扩展性问题。为此,本文开创性地将图遗忘与经典社会影响力最大化建立关联,提出可扩展的节点影响力最大化(NIM)方法,通过解耦传播模型与细粒度影响函数,实现无需依赖具体遗忘方法的离线执行,可无缝嵌入多数现有图遗忘框架以提升性能。基于此,我们进一步构建了可扩展图遗忘(SGU)新范式,通过实体特异性优化,在遗忘与推理能力间取得平衡。在14个数据集(含大规模ogbn-papers100M)上的实验表明,NIM显著提升多数现有方法的遗忘能力,而SGU达到全面领先性能并保持良好可扩展性。

原文摘要 · Abstract (English)

Machine unlearning, as a pivotal technology for enhancing model robustness and data privacy, has garnered significant attention in prevalent web mining applications, especially in thriving graph-based scenarios. However, most existing graph unlearning (GU) approaches face significant challenges due to the intricate interactions among web-scale graph elements during the model training: (1) The gradient-driven node entanglement hinders the complete knowledge removal in response to unlearning requests; (2) The billion-level graph elements in the web scenarios present inevitable scalability issues. To break the above limitations, we open up a new perspective by drawing a connection between GU and conventional social influence maximization. To this end, we propose Node Influence Maximization (NIM) through the decoupled influence propagation model and fine-grained influence function in a scalable manner, which is crafted to be a plug-and-play strategy to identify potential nodes affected by unlearning entities. This approach enables offline execution independent of GU, allowing it to be seamlessly integrated into most GU methods to improve their unlearning performance. Based on this, we introduce Scalable Graph Unlearning (SGU) as a new fine-tuned framework, which balances the forgetting and reasoning capability of the unlearned model by entity-specific optimizations. Extensive experiments on 14 datasets, including large-scale ogbn-papers100M, have demonstrated the effectiveness of our approach. Specifically, NIM enhances the forgetting capability of most GU methods, while SGU achieves comprehensive SOTA performance and maintains scalability.

图学习机器遗忘可扩展性隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。