提出多任务评测框架,评估大模型删去敏感内容的效果。
LUME: LLM Unlearning with Multitask Evaluations
- 构建三个真实场景的删忆任务:虚构小说、敏感人物传记、公开人物资料
- 发布10亿与70亿参数的微调模型作为目标,用于评测删忆效果
- 首次系统评测多种删忆算法在不同任务中的表现与局限
删忆旨在不进行完整重训练的情况下,从大语言模型中移除版权、敏感或隐私内容。本文构建了一个多任务删忆评测基准(LUME),包含三项任务:(1) 删去合成生成的创意短篇小说,(2) 删去含敏感信息的合成人物传记,(3) 删去一组公开人物传记。我们还发布了两个经过微调的1B和7B参数量的LLM作为目标模型。对近期提出的多种删忆算法进行了详尽评估,并采用精心设计的指标分析其行为与局限性。
原文摘要 · Abstract (English)
Unlearning aims to remove copyrighted, sensitive, or private content from large language models (LLMs) without a full retraining. In this work, we develop a multi-task unlearning benchmark (LUME) which features three tasks: (1) unlearn synthetically generated creative short novels, (2) unlearn synthetic biographies with sensitive information, and (3) unlearn a collection of public biographies. We further release two fine-tuned LLMs of 1B and 7B parameter sizes as the target models. We conduct detailed evaluations of several recently proposed unlearning algorithms and present results on carefully crafted metrics to understand their behavior and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。