让多模态大模型安全删记忆,还能保持回答质量。
ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

- 通过激活引导让模型主动拒绝输出敏感信息
- 在删记忆时提升生成质量5.8倍,效果提高24.6%
- 仅需少量数据就能实现安全且有用的去记忆
多模态大语言模型(MLLM)在预训练中可能记住敏感的跨模态信息,因此机器去记忆(MU)至关重要。现有方法通常以输出偏差评估去记忆效果,却忽视了去记忆后的生成质量,易导致幻觉或僵化回复,影响模型可用性与安全性。为此,我们提出ASRU,一种可控的多模态去记忆框架,将生成质量作为核心评估目标。ASRU首先通过激活重定向诱导初始拒绝行为,再利用定制奖励函数优化细粒度拒绝边界,从而在目标知识删除与模型实用性之间取得更好平衡。在Qwen3-VL上的实验表明,ASRU平均提升去记忆效果24.6%,生成质量提升5.8倍,同时有效保留模型能力,仅需少量保留监督数据即可实现。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) may memorize sensitive cross-modal information during pretraining, making machine unlearning (MU) crucial. Existing methods typically evaluate unlearning effectiveness based on output deviations, while overlooking the generation quality after unlearning. This can easily lead to hallucinated or rigid responses, thereby affecting the usability and safety of the unlearned model. To address this issue, we propose ASRU, a controllable multimodal unlearning framework that incorporates generation quality as a core evaluation objective. ASRU first induces initial refusal behavior through activation redirection, and then optimizes fine-grained refusal boundaries using a customized reward function, thereby achieving a better trade-off between target knowledge unlearning and model utility. Experiments on Qwen3-VL show that ASRU significantly improves unlearning effectiveness (+24.6%) on average and generation quality (5.8X) on average while effectively preserving model utility, using only a small amount of retained supervision data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。