arXiv:2504.02883cs.CLcs.LG2025-04ACL被引 8

评测大模型删除敏感内容的三种实用方法,助力安全可控的AI应用。

SemEval-2025 Task 4: Unlearning sensitive content from Large Language Models

  • 针对长篇创作、短篇个人资料和真实训练数据三类场景设计去敏任务
  • 超100份方案来自30多所机构,验证了多种有效去敏技术路径
  • 为实际部署中隐私保护与内容安全提供可落地的参考方案

我们介绍 SemEval-2025 Task 4:从大语言模型(LLMs)中去学习敏感内容。该任务涵盖三个子任务,覆盖不同应用场景:(1) 去除长篇合成创作文本(涉及多种文体);(2) 去除包含个人身份信息(PII)的短篇合成传记,包括虚假姓名、电话号码、社会安全号码(SSN)、电子邮件和家庭地址;(3) 去除从目标模型训练数据集中采样的真实文档。本次任务共收到来自30多个机构的100余份提交,本文总结了关键技术与核心经验教训。

原文摘要 · Abstract (English)

We introduce SemEval-2025 Task 4: unlearning sensitive content from Large Language Models (LLMs). The task features 3 subtasks for LLM unlearning spanning different use cases: (1) unlearn long form synthetic creative documents spanning different genres; (2) unlearn short form synthetic biographies containing personally identifiable information (PII), including fake names, phone number, SSN, email and home addresses, and (3) unlearn real documents sampled from the target model's training dataset. We received over 100 submissions from over 30 institutions and we summarize the key techniques and lessons in this paper.

大模型安全去敏隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。