用语音转换扩充德语方言数据,提升小样本分类效果
Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion
- 用检索式语音转换统一说话人,减少个体差异干扰
- 单独使用可提升分类准确率,与掩蔽等方法结合更优
- 适合数据稀缺的方言识别任务,尤其对小语种有效
针对方言识别中方言数据稀缺的问题,本文提出使用检索式语音转换(RVC)作为数据增强方法,用于低资源德语方言分类任务。通过将音频样本转换为统一目标说话人,RVC有效降低说话人相关变异性,使模型更聚焦于方言特有的语言和语音特征。实验表明,RVC作为独立增强手段即可提升分类性能;进一步与频率掩蔽、片段删除等方法结合,可带来额外增益,验证了其在低资源场景下改善方言分类的潜力。
原文摘要 · Abstract (English)
Deep learning models for dialect identification are often limited by the scarcity of dialectal data. To address this challenge, we propose to use Retrieval-based Voice Conversion (RVC) as an effective data augmentation method for a low-resource German dialect classification task. By converting audio samples to a uniform target speaker, RVC minimizes speaker-related variability, enabling models to focus on dialect-specific linguistic and phonetic features. Our experiments demonstrate that RVC enhances classification performance when utilized as a standalone augmentation method. Furthermore, combining RVC with other augmentation methods such as frequency masking and segment removal leads to additional performance gains, highlighting its potential for improving dialect classification in low-resource scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。