arXiv:2606.07141cs.LGcs.AI2026-06

构建医疗疾病推断的遗忘评估基准,解决隐私数据移除难题。

REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference

论文配图:REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference
图 1 · 摘自论文原文
  • 基于MIMIC-III数据库构建多标签医疗推理遗忘评估基准
  • 发现现有遗忘方法在多标签任务中性能下降且存在效用与遗忘的权衡
  • 适合医疗AI安全、隐私合规及模型可解释性研究者使用

针对临床疾病推断语言模型训练中包含敏感患者数据的问题,当数据所有者要求移除其数据时,精确遗忘个体数据难以实现,而微调重训又成本高昂。现有机器遗忘方法多适用于非医疗领域,且缺乏真实医疗场景下的评估基准。为此,我们提出REMEDI,一个面向多标签、多分类临床疾病推断的遗忘评估基准,涵盖标签相关性、纵向数据结构与安全约束等挑战。该基准基于包含完整临床信息的MIMIC-III数据库构建,支持多种遗忘实例设置,并采用兼顾效用与遗忘程度的综合评估指标。实验表明,现有方法在多标签任务中表现不佳,且存在效用与遗忘效果的权衡。为促进复现,基准已公开。

原文摘要 · Abstract (English)

Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owners may request the removal of their data from a trained model due to privacy or copyright concerns. However, exactly unlearning patient-specific data is intractable, and retraining with minor data removal is resource-intensive. While there exists several machine unlearning methods that can be used, their utility is generally restricted to non-medical domains. Moreover, the existing benchmarks for evaluating such unlearning methods primarily utilize synthetically curated datasets, which are not truly representative of real-world systems. Hence, the effectiveness of these unlearning methods in the medical domain is largely unclear. To this end, we introduce REMEDI, an extensive benchmark for machine unlearning tailored to multi-label and multiclass clinical disease inference, where label correlations, longitudinal structure, and safety constraints make unlearning particularly challenging. Unlike the existing benchmarks, REMEDI considers: (1) a relevant application domain (medical), (2) comprehensive unlearning setups involving diverse sets of forget instances, (3) challenging unlearning scenarios including multi-label and multi-class classification tasks, and (4) evaluation metrics involving performance both in terms of utility and extent of unlearning achieved. REMEDI is developed using the MIMIC-III clinical database that contains comprehensive clinical data of patients. Experiments with existing unlearning methods indicate that there exists a trade-off between utility and unlearning performance. They are also largely unsuited to multi-label classification tasks. To facilitate reproducibility, we make our benchmark publicly available.

医疗AI模型遗忘隐私保护多标签学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。