为大模型遗忘提供可控制的风险管理框架,平衡遗忘效果与模型性能。
FROC: A Unified Framework with Risk-Optimized Control for Machine Unlearning in LLMs
- 基于概率约束的统一框架,量化遗忘风险
- 通过连续风险模型评估遗忘不足与性能损失
- 帮助选择安全且高效的遗忘策略,适合合规场景
机器遗忘旨在消除特定训练样本对已部署模型的影响。随着大语言模型广泛应用,不充分遗忘或性能损失带来的风险日益突出。现有遗忘技术缺乏有效的风险评估与控制机制,难以在安全与性能间做出合理权衡,引发对“被遗忘权”的信任危机。为此,我们提出FROC——一种面向大语言模型机器遗忘的统一风险优化控制框架。FROC采用类似置信区间的风险控制范式,以用户指定的风险预算约束遗忘行为。该概率性约束使FROC能够比较不同遗忘策略、识别可行操作区域,并根据期望的遗忘充分性与性能保留之间的权衡指导超参数选择。为实现该约束,FROC引入平滑连续的风险模型,将遗忘不足与性能退化整合为单一配置级评分。基于置信风险分析,FROC计算(1)置信遗忘风险(CUR),即被遗忘样本仍影响模型预测的概率的非参数估计值;(2)风险可控配置集,识别在给定风险预算下有效的遗忘超参数。多个大语言模型遗忘方法的实验表明,FROC生成稳定可解释的风险图景,揭示了遗忘配置、语义偏移与性能影响间的规律性关系。FROC将机器遗忘重构为可控、风险感知的过程,为大规模大语言模型部署中的遗忘行为管理提供了实用基础。
原文摘要 · Abstract (English)
Machine unlearning (MU) seeks to eliminate the influence of specific training examples from deployed models. As large language models (LLMs) become widely used, managing risks arising from insufficient forgetting or utility loss is increasingly crucial. Current MU techniques lack effective mechanisms for evaluating and controlling these risks, hindering the selection of strategies that appropriately balance safety and utility, and raising trust concerns surrounding the "right to be forgotten." To address these issues, we propose FROC, a unified framework with Risk-Optimized Control for machine unlearning in LLMs. FROC is built around a conformal-style risk-control formulation that expresses a user-specified risk budget on unlearning behavior. This probability-based constraint enables FROC to compare MU strategies, identify feasible operating regions, and guide hyperparameter selection according to desired trade-offs between forgetting sufficiency and utility preservation. To operationalize this constraint, FROC introduces a smoothly varying continuous risk model that aggregates forgetting deficiency and utility degradation into a single configuration-level score. Building on conformal risk analysis, FROC computes (1) the Conformal Unlearning Risk (CUR), a data-driven estimated value on the probability that forgotten samples continue to influence model predictions, and (2) risk-controlled configuration sets, which identify unlearning hyperparameters that are valid under the specified risk budget. Experiments across multiple LLM MU methods demonstrate that FROC produces stable, interpretable risk landscapes and reveals consistent relationships between unlearning configurations, semantic shift, and utility impact. FROC reframes MU as a controllable, risk-aware process and offers a practical foundation for managing unlearning behavior in large-scale LLM deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。