arXiv:2412.08637cs.CVcs.AI2024-12中稿 · CVPR被引 5

提出可扩展的扩散模型数据影响估计方法,支持百亿参数模型

DMin: Scalable Training Data Influence Estimation for Diffusion Models

  • 基于高效梯度压缩,大幅降低存储需求
  • 在1秒内完成百万级样本中前k个关键样本检索
  • 首个支持百亿参数扩散模型的影响评估工具

识别对生成图像影响最大的训练数据样本是理解扩散模型的关键任务,但现有方法受限于计算成本,仅适用于小规模或LoRA微调模型。为此,我们提出DMin(Diffusion Model influence),首个可扩展的扩散模型训练数据影响估计框架,支持拥有数十亿参数的模型。通过高效梯度压缩技术,DMin将存储需求从数百TB降至MB甚至KB级别,并可在1秒内完成对前k个最具影响力的训练样本的检索,同时保持高精度。实验表明,DMin在识别关键训练样本方面有效,且在计算与存储效率上表现卓越。

原文摘要 · Abstract (English)

Identifying the training data samples that most influence a generated image is a critical task in understanding diffusion models (DMs), yet existing influence estimation methods are constrained to small-scale or LoRA-tuned models due to computational limitations. To address this challenge, we propose DMin (Diffusion Model influence), a scalable framework for estimating the influence of each training data sample on a given generated image. To the best of our knowledge, it is the first method capable of influence estimation for DMs with billions of parameters. Leveraging efficient gradient compression, DMin reduces storage requirements from hundreds of TBs to mere MBs or even KBs, and retrieves the top-k most influential training samples in under 1 second, all while maintaining performance. Our empirical results demonstrate DMin is both effective in identifying influential training samples and efficient in terms of computational and storage requirements.

扩散模型数据影响高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。