arXiv:2412.10649cs.SDcs.AI2024-12

在音频生成模型中隐藏不可察觉的回声,可有效追踪训练数据来源。

Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models

  • 在训练数据中嵌入微弱回声,各类音频生成模型均能复现
  • 单个回声在多种架构中稳定存在,长时序回声提升信息容量
  • 回声可抵抗微调、混音、变调等处理,适合用于模型水印

随着生成式技术在音频领域的普及,人们愈发关注如何追溯这些复杂模型对训练数据的依赖关系,以确保数据使用合规并揭示其黑箱行为。本文发现,若训练数据中隐藏了不可察觉的回声,多种音频到音频生成架构(包括可微数字信号处理(DDSP)、实时音频变分自编码器(RAVE)及“Dance Diffusion”)均会在输出中重现这些回声。单个回声在所有架构中均表现稳健;我们还展示了更长时序分布的回声模式在提升信息容量方面具有潜力。进一步实验表明,回声可存在于微调后的模型中,且在混合/分离和训练阶段的音高变换增强下依然存活。这表明经典的水印思路在生成音频模型中具有显著应用前景。

原文摘要 · Abstract (English)

As generative techniques pervade the audio domain, there has been increasing interest in tracing back through these complicated models to understand how they draw on their training data to synthesize new examples, both to ensure that they use properly licensed data and also to elucidate their black box behavior. In this paper, we show that if imperceptible echoes are hidden in the training data, a wide variety of audio to audio architectures (differentiable digital signal processing (DDSP), Realtime Audio Variational autoEncoder (RAVE), and ``Dance Diffusion'') will reproduce these echoes in their outputs. Hiding a single echo is particularly robust across all architectures, but we also show promising results hiding longer time spread echo patterns for an increased information capacity. We conclude by showing that echoes make their way into fine tuned models, that they survive mixing/demixing, and that they survive pitch shift augmentation during training. Hence, this simple, classical idea in watermarking shows significant promise for tagging generative audio models.

音频生成模型水印数据溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。