arXiv:2506.18312cs.SDeess.AS2025-06中稿 · NeurIPS被引 10

用反学习技术追踪音乐生成模型的训练数据来源,实现版权溯源。

Large-Scale Training Data Attribution for Music Generative Models via Unlearning

  • 通过反学习方法识别生成音乐时最关键的训练数据点。
  • 在大规模文本到音乐扩散模型上验证了方法的有效性与一致性。
  • 为音乐AI版权归属提供可解释的追溯方案,适合伦理与法律研究者。

本文探索在大规模数据集训练的音乐生成模型中,利用反学习方法进行训练数据溯源(TDA)。TDA旨在识别特定模型输出中贡献最大的训练数据点,这对人工智能生成音乐中的原创艺术家署名问题至关重要。通过白盒溯源,本工作支持更公平的艺术贡献认可,回应人工智能伦理与版权的紧迫关切。我们将在大规模文本到音乐扩散模型上应用反学习溯源,并通过网格搜索不同超参数配置,定量评估反学习方法的一致性。随后将反学习的溯源模式与非反事实方法进行对比。结果表明,反学习方法可有效适配音乐生成模型,首次引入大规模音乐生成领域的训练数据溯源,为更伦理、可问责的音乐生成AI系统铺平道路。

原文摘要 · Abstract (English)

This paper explores the use of unlearning methods for training data attribution (TDA) in music generative models trained on large-scale datasets. TDA aims to identify which specific training data points contributed the most to the generation of a particular output from a specific model. This is crucial in the context of AI-generated music, where proper recognition and credit for original artists are generally overlooked. By enabling white-box attribution, our work supports a fairer system for acknowledging artistic contributions and addresses pressing concerns related to AI ethics and copyright. We apply unlearning-based attribution to a text-to-music diffusion model trained on a large-scale dataset and investigate its feasibility and behavior in this setting. To validate the method, we perform a grid search over different hyperparameter configurations and quantitatively evaluate the consistency of the unlearning approach. We then compare attribution patterns from unlearning with non-counterfactual approaches. Our findings suggest that unlearning-based approaches can be effectively adapted to music generative models, introducing large-scale TDA to this domain and paving the way for more ethical and accountable AI systems for music creation.

音乐生成数据溯源反学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。