arXiv:2607.13037cs.AI2026-07

精准定位数据作者贡献,实现删除请求的细粒度处理。

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

  • 通过记录与标记级溯源,追踪数据来源至具体条目和词元。
  • 减少98.7%的过度删除,仅需1.3-19.0%额外计算开销。
  • 适合需合规删除、隐私保护或模型可解释性的研究者使用。

当数据贡献者申请删除时,模型训练方面临实际困境:遗忘算法需要精确的删除集合,但现有工具无法定位特定作者的数据记录。现有溯源系统仅支持文件或数据集级别,导致灾难性过量删除。本文提出ob,一种记录与词元级别的数据溯源系统,可在数据处理流程中传播作者身份,并通过确定性查询将删除请求转化为精确的遗忘集合。在219,555篇维基百科页面上的评估显示,记录级溯源将过量删除从101倍降至1.3倍;集成后在HuggingFace上增加1.3-4.0%吞吐开销,在Datatrove上为2.1-19.0%(维基数据)。在17亿参数模型上,基于溯源的遗忘集合在所有被评估作者中均显著降低机器遗忘的附带损害(保持困惑度),且成员推理测试表明其选取内容确为模型记忆部分。

原文摘要 · Abstract (English)

When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author. Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation requests into precise forget sets via deterministic queries. Evaluation on 219,555 Wikipedia pages demonstrates that record-level provenance eliminates dataset-level over-deletion (from 101x to 1.3x), while integration adds 1.3-4.0% throughput overhead (HuggingFace) and 2.1-19.0% (Datatrove) on wiki data. On a 1.7B model, provenance-based forget sets consistently reduce the collateral damage of machine unlearning (retain perplexity) relative to same-size random baselines across all evaluated authors, with membership-inference tests indicating they select genuinely memorized content.

数据溯源隐私保护模型遗忘维基数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。