针对多模态大模型推理泄露隐私问题,提出无需训练的实时净化方法。
LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

- 利用推理过程中的熵动态识别敏感内容起始点
- 通过视觉锚点注入替换敏感词元,抑制推理与答案泄露
- 适用于具备强化学习推理能力的多模态模型,保护隐私且不损失通用性能
强化学习后训练使多模态大推理模型(MLRMs)具备探索性思维链(CoT),显著提升视觉推理能力。然而我们发现,这一能力带来独特隐私漏洞:即使最终答案已成功清除敏感信息,推理过程仍可能重现该信息。这种泄漏在原生强化学习训练的模型中远高于非推理基线模型,现有去学习方法无法应对。我们观察到,强化学习引发的探索会在敏感内容上留下独特的词元级熵特征,而基线模型中几乎不存在。基于此,提出 LEMUR——一种完全无训练、推理时生效的去学习框架。LEMUR以熵动态为控制信号,精准定位敏感推理开始与终止时刻,在此区间内通过熵调节的视觉锚点潜变量注入,将已确定的词元替换为重新锚定至输入图像的净化嵌入。在多种 MLRM 上,LEMUR 均显著优于现有去学习方法,在抑制推理轨迹与答案泄露的同时,更好保持非敏感任务性能与输出流畅性。结果表明,强化学习诱导的熵动态是隐私泄露的独特信号,利用该信号可实现高效无训练的推理型多模态模型去学习。
原文摘要 · Abstract (English)
Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non -reasoning base models, revealing a privacy risk that existing unlearning methods are not designed to address. We show that RL-induced exploration leaves sensitive content with a distinctive token-level entropy signature that is largely absent from base models. Based on this observation, we propose LEMUR, a fully training-free, inference-time unlearning framework for natively RL-trained multimodal models. LEMUR uses entropy dynamics as a control signal to identify when sensitive reasoning begins and when sanitization should stop. During this interval, it redirects the reasoning trajectory through entropy-modulated visual-anchor latent injection, replacing committed tokens with sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse MLRMs, LEMUR consistently outperforms existing unlearning met hods in suppressing both reasoning-trace and answer leakage, while better preserving non-sensitive utility and output fluency. These results demonstrate that RL-induced entropy dynamics provide a distinctive signal for privacy leakage and that exploiting this signal enables effective training-free unlearning for reasoning-capable multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。