arXiv:2607.21300cs.CVcs.AI2026-07

针对多模态大模型遗忘中的不公平问题,提出首个真实不平衡场景下的评测与算法。

Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning

论文配图:Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning
图 1 · 摘自论文原文
  • 构建真实分布的遗忘请求基准FAIRGET,模拟不同群体遗忘频率差异
  • 提出公平感知的遗忘算法FAUN,避免模型对特定群体产生偏见
  • 在多个数据集上验证,同时提升遗忘效果与模型公平性

机器遗忘作为满足最新人工智能监管要求的重要工具,用于从训练好的模型中移除个人数据。现有研究在多模态大语言模型(MLLMs)上评估遗忘效果时,通常通过微调虚构身份来模拟遗忘请求,且这些身份在数据中均匀分布。然而,在现实场景中,不同人口群体提出遗忘请求的频率可能存在差异,可能改变模型对这些群体的内部认知,导致偏见行为。为填补这一空白,我们提出FAIRGET,首个在非均衡、真实遗忘请求下评估多模态大模型遗忘效果的视觉问答基准。该基准设计了从简单到复杂的多种真实场景,若不考虑公平性,将导致遗忘后模型产生偏见。此外,我们提出首个面向MLLMs的遗忘算法FAUN,能够在遗忘特定数据的同时保持模型公平性。FAUN利用一种偏见感知的激活引导机制,以应对遗忘数据的非均衡特性。在FAIRGET和现有基准FIUBench上的实验表明,本方法在遗忘质量和公平性方面均优于现有方法。

原文摘要 · Abstract (English)

Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning effectiveness in multimodal large language models (MLLMs), prior works fine-tune models on fictitious identities, simulating unlearning requests on subsets of these IDs, which are typically uniformly distributed. However, in realistic scenarios, people from different demographic groups may request to be unlearned at different frequencies, potentially altering the model's internal beliefs for these groups and leading to biased behaviors. To fill this gap, we propose FAIRGET, the first Visual Question Answering benchmark that evaluates unlearning under unbalanced, realistic, forget requests. These requests are designed to simulate multiple realistic scenarios, ranging from simple to challenging settings, that lead to biased unlearned models if fairness is not accounted for. Additionally, we propose FAUN, the first unlearning algorithm for MLLMs that forgets unlearning data while preserving model fairness. FAUN exploits a bias-aware activation steering mechanism to unlearn identities while accounting for the unbalanced nature of the forget data. Experiments on FAIRGET and the established FIUBench demonstrate our method's superiority both in unlearning quality and fairness.

多模态遗忘学习公平性评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。