用定制提示生成音频数据,提升声学分类效果
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
- 设计任务导向的提示策略生成更真实音频
- 融合多模型生成数据比单纯扩容更有效
- 适合需要合成数据增强的声学研究者
本文研究了利用文本转音频(TTA)模型生成真实感数据集时的有效提示策略,并分析了不同方法组合这些数据集以提升声学分类任务性能的效果。通过在两个声学分类数据集上使用两种TTA模型,我们测试了多种提示策略。结果表明,任务特定的提示策略在数据生成上显著优于基础提示方法。此外,使用不同TTA模型生成的数据集进行合并,比单纯增加训练数据规模更能有效提升分类表现。总体而言,这些方法作为基于合成数据的数据增强技术具有明显优势。
原文摘要 · Abstract (English)
This paper investigates the design of effective prompt strategies for generating realistic datasets using Text-To-Audio (TTA) models. We also analyze different techniques for efficiently combining these datasets to enhance their utility in sound classification tasks. By evaluating two sound classification datasets with two TTA models, we apply a range of prompt strategies. Our findings reveal that task-specific prompt strategies significantly outperform basic prompt approaches in data generation. Furthermore, merging datasets generated using different TTA models proves to enhance classification performance more effectively than merely increasing the training dataset size. Overall, our results underscore the advantages of these methods as effective data augmentation techniques using synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。