arXiv:2412.08549cs.LGcs.SD2024-12被引 7

用音频水印追踪音乐生成模型的训练数据来源

Watermarking Training Data of Music Generation Models

  • 在训练数据中嵌入人耳不可察觉的音频水印
  • 水印导致生成音乐出现可检测的风格偏移,比例越高越明显
  • 适合关注版权保护与模型溯源的研究者

生成式人工智能模型在文本、图像和音频等领域广泛应用。这些模型的能力源自对海量人类创作内容的训练,其中常包含受版权保护的材料。本文研究音频水印技术能否用于检测音乐生成模型未经许可使用训练数据的情况。通过对比在含水印数据和无水印数据上训练的模型输出,分析水印技术、水印数据占比以及水印对模型分词器的鲁棒性等因素对生成行为的影响。结果表明,包括人耳不可察觉的水印技术在内,均会导致模型输出产生显著变化。同时评估了先进水印技术对移除攻击的鲁棒性。

原文摘要 · Abstract (English)

Generative Artificial Intelligence (Gen-AI) models are increasingly used to produce content across domains, including text, images, and audio. While these models represent a major technical breakthrough, they gain their generative capabilities from being trained on enormous amounts of human-generated content, which often includes copyrighted material. In this work, we investigate whether audio watermarking techniques can be used to detect an unauthorized usage of content to train a music generation model. We compare outputs generated by a model trained on watermarked data to a model trained on non-watermarked data. We study factors that impact the model's generation behaviour: the watermarking technique, the proportion of watermarked samples in the training set, and the robustness of the watermarking technique against the model's tokenizer. Our results show that audio watermarking techniques, including some that are imperceptible to humans, can lead to noticeable shifts in the model's outputs. We also study the robustness of a state-of-the-art watermarking technique to removal techniques.

音频生成水印技术版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。