arXiv:2509.14934eess.AScs.LG2025-09中稿 · ICASSP 2026被引 2

用反记忆引导减少文本到音频模型的数据复现问题

Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance

  • 在采样过程中引入反记忆引导,抑制模型复现训练数据
  • 三种引导策略有效降低复现率,同时保持音质和语义一致性
  • 适用于注重生成安全性的音频生成系统开发者

生成式音频模型普遍存在数据复现问题,即模型在推理时无意生成训练数据中的部分内容。本文针对文本到音频扩散模型,探索使用反记忆策略缓解该问题。采用反记忆引导(AMG)技术,通过修改预训练扩散模型的采样过程,抑制记忆行为。研究评估了三种不同类型的AMG引导策略,均旨在减少复现现象的同时保持生成质量。以Stable Audio Open为基线模型,利用其完全开源的架构与训练数据集进行实验。全面的分析表明,AMG显著降低了扩散模型在文本到音频生成中的记忆复现问题,且未损害音频保真度或语义对齐性能。

原文摘要 · Abstract (English)

A persistent challenge in generative audio models is data replication, where the model unintentionally generates parts of its training data during inference. In this work, we address this issue in text-to-audio diffusion models by exploring the use of anti-memorization strategies. We adopt Anti-Memorization Guidance (AMG), a technique that modifies the sampling process of pre-trained diffusion models to discourage memorization. Our study explores three types of guidance within AMG, each designed to reduce replication while preserving generation quality. We use Stable Audio Open as our backbone, leveraging its fully open-source architecture and training dataset. Our comprehensive experimental analysis suggests that AMG significantly mitigates memorization in diffusion-based text-to-audio generation without compromising audio fidelity or semantic alignment.

音频生成扩散模型生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。