构建多语言虚构演员数据集,评估大模型删除特定信息的能力
FAME: Fictional Actors for Multilingual Erasure
- 用虚构人物构建多语言数据集,支持实体与实例级遗忘
- 覆盖5种语言,含1000条人物传记和2万问答对,每条含20个主题
- 专为隐私保护设计,确保训练时未接触过这些数据
基于网络规模数据训练的大语言模型引发隐私与被遗忘权担忧。机器遗忘技术可在不重新训练的前提下移除模型中的特定信息。但现有大模型遗忘评估基准存在两大局限:仅支持英文,且仅能实现实体级遗忘(删除某人全部信息)。本文提出FAME(Fictional Actors for Multilingual Erasure),一个涵盖英语、法语、德语、意大利语和西班牙语的合成基准,包含1,000条虚构演员传记和20,000个问答对。每条传记覆盖20个主题,分属人物背景、职业生涯、成就和个人信息等结构化类别,支持实体级遗忘(删除完整身份)与实例级遗忘(删除特定事实而保留其他信息)。提供两个数据集划分以支持不同遗忘场景,并实现跨语言方法的系统性对比。由于数据完全虚构,未在模型预训练中出现,可实现对遗忘方法的可控评估。
原文摘要 · Abstract (English)
LLMs trained on web-scale data raise concerns about privacy and the right to be forgotten. To address these issues, Machine Unlearning provides techniques to remove specific information from trained models without retraining from scratch. However, existing benchmarks for evaluating unlearning in LLMs face two major limitations: they focus only on English and support only entity-level forgetting (removing all information about a person). We introduce FAME (Fictional Actors for Multilingual Erasure), a synthetic benchmark for evaluating Machine Unlearning across five languages: English, French, German, Italian, and Spanish. FAME contains 1,000 fictional actor biographies and 20,000 question-answer pairs. Each biography includes information on 20 topics organized into structured categories (biography, career, achievements, personal information). This design enables both entity-level unlearning (i.e., forgetting entire identities) and instance-level unlearning (i.e., forgetting specific facts while retaining others). We provide two dataset splits to support these two different unlearning scenarios and enable systematic comparison of unlearning techniques across languages. Since FAME uses entirely fictional data, it ensures that the information was never encountered during model pretraining, allowing for a controlled evaluation of unlearning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。