用自适应方法从大模型中清除敏感信息,提升隐私安全性。
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
- 采用自适应表示误导机制,针对性移除模型中的敏感知识。
- 在1B和7B参数模型上均取得第4名,验证了方法有效性。
- 适合关注模型隐私保护与数据合规的研究者使用。
大型语言模型(LLMs)在自然语言理解与生成方面表现出色,但其对训练数据的过度记忆引发了隐私、版权合规及安全问题,尤其是在涉及个人身份信息(PII)时。有效的机器遗忘技术对缓解此类风险至关重要,然而现有方法在面向LLMs时仍不成熟,主要受限于其开放输出空间。本文应用自适应表示误导遗忘(Adaptive RMU)技术,对LLMs中的敏感信息进行清除。通过大量实验,分析了不同解码器层在遗忘过程中的效果,确定了最适宜敏感信息移除的区域。该方法在SemEval-2025任务4的官方排行榜上,分别在1B参数和7B参数模型中位列第4名。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation. However, their tendency to memorize training data raises concerns regarding privacy, copyright compliance, and security, particularly in cases involving Personally Identifiable Information (PII). Effective machine unlearning techniques are essential to mitigate these risks, yet existing methods remain underdeveloped for LLMs due to their open-ended output space. In this work, we apply the Adaptive Representation Misdirection Unlearning (RMU) technique to unlearn sensitive information from LLMs. Through extensive experiments, we analyze the effects of unlearning across different decoder layers to determine the most effective regions for sensitive information removal. Our technique ranked 4th on the official leaderboard of both 1B parameter and 7B parameter models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。