arXiv:2511.03966cs.LG2025-11被引 1

提出分层遗忘算法HIF,高效删除认知诊断模型中的学生数据

PrivacyCD: Hierarchical Unlearning for Protecting Student Privacy in Cognitive Diagnosis

  • 基于参数重要性分层特性,设计混合重要度平滑机制
  • 在3个真实数据集上显著优于基线,平衡遗忘效果与模型性能
  • 适合需要响应数据删除请求的教育AI系统开发者

随着用户对“被遗忘权”的日益重视,从认知诊断(CD)模型中移除特定学生数据已成为迫切需求。然而现有CD模型普遍缺乏隐私设计,且缺少有效的数据遗忘机制。直接套用通用遗忘算法效果不佳,难以兼顾遗忘完整性、模型效用与效率,尤其在处理CD模型特有的异构结构时表现欠佳。本文首次系统研究了CD模型的数据遗忘问题,提出一种新颖高效的算法:分层重要性引导遗忘(HIF)。核心洞察是CD模型中参数重要性具有明显的分层特征。HIF通过创新的平滑机制,融合个体与层级层面的重要性,更精准识别与待删除数据相关的参数。在三个真实世界数据集上的实验表明,HIF在关键指标上显著优于基线,为CD模型应对用户数据删除请求提供了首个有效解决方案,助力部署高性能、隐私保护的AI系统。

原文摘要 · Abstract (English)

The need to remove specific student data from cognitive diagnosis (CD) models has become a pressing requirement, driven by users' growing assertion of their "right to be forgotten". However, existing CD models are largely designed without privacy considerations and lack effective data unlearning mechanisms. Directly applying general purpose unlearning algorithms is suboptimal, as they struggle to balance unlearning completeness, model utility, and efficiency when confronted with the unique heterogeneous structure of CD models. To address this, our paper presents the first systematic study of the data unlearning problem for CD models, proposing a novel and efficient algorithm: hierarchical importanceguided forgetting (HIF). Our key insight is that parameter importance in CD models exhibits distinct layer wise characteristics. HIF leverages this via an innovative smoothing mechanism that combines individual and layer, level importance, enabling a more precise distinction of parameters associated with the data to be unlearned. Experiments on three real world datasets show that HIF significantly outperforms baselines on key metrics, offering the first effective solution for CD models to respond to user data removal requests and for deploying high-performance, privacy preserving AI systems

认知诊断数据遗忘隐私保护分层机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。