提出HANKER框架,自动生成大规模去记忆审计数据集,提升评估准确性。
Holistic Audit Dataset Generation for LLM Unlearning via Knowledge Graph Traversal and Redundancy Removal
- 基于知识图谱遍历与去重,实现细粒度覆盖的审计数据生成
- 在新闻和书籍数据集上分别生成超6.9万和11.1万个测试用例
- 揭示冗余信息导致评估结果虚高,强调去重对真实评估的重要性
近年来,大语言模型需通过机器去记忆来选择性删除敏感信息、保护隐私并遵守版权法规。然而,现有评估基准规模与全面性不足,通常仅包含数百个测试案例。我们识别出两大挑战:确保审计充分性及处理遗忘与保留数据集间的知识冗余。为此,提出HANKER自动化框架,利用知识图谱实现细粒度覆盖并消除冗余知识。将HANKER应用于MUSE基准,成功生成新闻与书籍数据集分别超过69,000和111,000个审计案例,发现数千例原基准未能检测到的知识记忆实例。实证分析显示,知识冗余显著扭曲去记忆评估指标:ROUGE从19.7%升至26.1%,蕴含度分数从32.4%增至35.2%,凸显系统性去重对准确评估的必要性。
原文摘要 · Abstract (English)
In recent years, Large Language Models (LLMs) have faced increasing demands to selectively remove sensitive information, protect privacy, and comply with copyright regulations through unlearning, by Machine Unlearning. While evaluating unlearning effectiveness is crucial, existing benchmarks are limited in scale and comprehensiveness, typically containing only a few hundred test cases. We identify two critical challenges in generating holistic audit datasets: ensuring audit adequacy and handling knowledge redundancy between forget and retain dataset. To address these challenges, we propose HANKER, an automated framework for holistic audit dataset generation leveraging knowledge graphs to achieve fine-grained coverage and eliminate redundant knowledge. Applying HANKER to the popular MUSE benchmark, we successfully generated over 69,000 and 111,000 audit cases for the News and Books datasets respectively, identifying thousands of knowledge memorization instances that the previous benchmark failed to detect. Our empirical analysis uncovers how knowledge redundancy significantly skews unlearning effectiveness metrics, with redundant instances artificially inflating the observed memorization measurements ROUGE from 19.7% to 26.1% and Entailment Scores from 32.4% to 35.2%, highlighting the necessity of systematic deduplication for accurate assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。