用户给参考音频,模型就能生成包含特定声音的定制化音频。
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- 通过参考音频提取个性化声音特征,实现细粒度控制
- 生成音频与文本提示高度一致,且保留指定声音特征
- 提供真实场景数据集,推动定制化语音生成研究
随着基于大规模扩散模型和语言模型的生成技术发展,文本到音频生成已取得显著进展。然而,现有模型主要关注语义对齐,难以精确控制特定声音的细微声学特性。为解决此问题,本文提出DreamAudio,一种定制化文本到音频生成(CTTA)框架。该框架能从用户提供的参考概念中识别听觉信息,仅需少量包含个性化音频事件的参考样本,即可生成包含这些特定事件的新音频。同时,我们构建了两类数据集用于训练与测试。实验表明,DreamAudio生成的音频在保持与文本提示高度一致的同时,能有效保留定制化声学特征。此外,其在通用文本到音频任务中表现相当。我们还提供了包含真实世界定制化案例的真人参与数据集,作为该任务的基准评测集。
原文摘要 · Abstract (English)
With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models mainly aim to generate semantically aligned sound and fall short of controlling fine-grained acoustic characteristics of specific sounds. As a result, users who need specific sound content may find it difficult to generate the desired audio clips. In this paper, we present DreamAudio for customized text-to-audio generation (CTTA). Specifically, we introduce a new framework that is designed to enable the model to identify auditory information from user-provided reference concepts for audio generation. Given a few reference audio samples containing personalized audio events, our system can generate new audio samples that include these specific events. In addition, two types of datasets are developed for training and testing the proposed systems. The experiments show that DreamAudio generates audio samples that are highly consistent with the customized audio features and aligned well with the input text prompts. Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。