arXiv:2604.13777cs.CLcs.AI2026-04被引 1

仅凭一个锚点即可实现大模型无须原始数据的精准擦除。

From Anchors to Supervision: Memory-Graph Guided Corpus-Free Unlearning for Large Language Models

论文配图:From Anchors to Supervision: Memory-Graph Guided Corpus-Free Unlearning for Large Language Models
图 1 · 摘自论文原文
  • 用轻量锚点探测模型记忆,构建加权记忆图。
  • 自动生成监督信号,效果接近外部参考数据。
  • 无需训练数据,适合隐私敏感场景使用。

大语言模型可能记忆敏感或受版权保护的内容,引发重大隐私与法律风险。尽管机器遗忘已成潜在解决方案,但现有方法依赖用户提供的遗忘集,导致遗忘请求难以审计,且存在二次泄露和恶意滥用风险。本文提出MAGE框架,一种基于记忆图引导的用户最小化、无需语料库的遗忘方法。仅需一个标识目标实体的轻量级用户锚点,MAGE即可探测目标模型以恢复相关记忆,将其组织为加权局部记忆图,并合成针对性的监督信号用于遗忘。MAGE具有模型无关性,可无缝集成至标准遗忘方法中,且无需访问原始训练语料。在两个基准数据集TOFU和RWKU上的实验表明,MAGE自生成的监督信号实现了与外部参考生成监督相当的遗忘效果,同时保持了模型整体性能。结果支持一种由极简锚点驱动、可审计的实用遗忘工作流,取代用户提交遗忘语料的传统模式。

原文摘要 · Abstract (English)

Large language models (LLMs) may memorize sensitive or copyrighted content, raising significant privacy and legal concerns. While machine unlearning has emerged as a potential remedy, prevailing paradigms rely on user-provided forget sets, making unlearning requests difficult to audit and exposing systems to secondary leakage and malicious abuse. We propose MAGE, a Memory-grAph Guided Erasure framework for user-minimized, corpus-free unlearning. Given only a lightweight user anchor that identifies a target entity, MAGE probes the target LLM to recover target-related memorization, organizes it into a weighted local memory graph, and synthesizes scoped supervision for unlearning. MAGE is model-agnostic, can be plugged into standard unlearning methods, and requires no access to the original training corpus. Experiments on two benchmarks, TOFU and RWKU, demonstrate that MAGE's self-generated supervision achieves effective unlearning performance comparable to supervision generated with external reference, while preserving overall utility. These results support a practical and auditable unlearning workflow driven by minimal anchors rather than user-supplied forget corpora.

大模型遗忘隐私保护记忆管理无语料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。