arXiv:2605.17938cs.LGcs.AI2026-05

提出MUCS方法,实现扩散模型训练数据溯源的高可靠性和鲁棒性。

Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew

论文配图:Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
图 1 · 摘自论文原文
  • 通过镜像反向梯度微调和噪声一致偏斜度衡量,实现数据溯源。
  • 在三个数据集上显著优于现有方法,性能提升明显。
  • 适合需要数据溯源与模型可解释性的研究者使用。

训练数据溯源(TDA)应促进生成模型的可解释性,并支持多种下游任务。然而,当前TDA方法缺乏可靠性与鲁棒性,难以应用于真实场景。本文提出一种基于镜像去学习与噪声一致偏斜度(MUCS)的TDA方法,通过有界镜像梯度上升微调第二模型,并利用一致噪声样本测量其相对于原模型的归一化偏斜度。实验表明,尽管方法概念简单且通用,MUCS在三个不同数据集上均显著优于现有方法。我们还分析了核心设计选择对性能的影响,探讨了生成样本中关键实例重叠特性及集成TDA方法的潜力。研究结果或对更广泛的去学习场景及扩散损失比较任务具有深远意义。

原文摘要 · Abstract (English)

Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliability and robustness, preventing their adoption in real-world setups. In this paper, we take a decisive step towards more reliable and robust TDA for diffusion models. We propose to perform TDA with mirrored unlearning and noise-consistent skew (MUCS). The idea is to fine-tune a second model with bounded mirrored gradient ascent, and to measure the normalized skew of this model with respect to the original one using consistent noise samples. We show that, while being conceptually simple and generic, MUCS systematically outperforms existing methods on three different datasets by a large margin. We additionally study the effect that core design choices have on final performance, and analyze novel aspects regarding the overlap of influential instances across generated items and the potential of ensembling TDA approaches. We believe that our findings may have broader implications for more general unlearning setups, as well as for tasks requiring the comparison of diffusion losses.

扩散模型数据溯源可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。