用更少计算资源实现文本无关语音模型的快速训练与高质量音频重建
Textless NLP -- Zero Resource Challenge with Low Resource Compute
- 采用优化的步长、插值因子和周期学习率调度加速收敛
- 在英、泰米尔、孟加拉语数据上均实现清晰语音重建,训练时间大幅缩短
- 适合低资源环境下语音合成与声纹转换研究者参考
本文针对文本无关自然语言处理中轻量级编码器-声码器模型仍需大量训练时间和GPU资源的问题,提出改进方案:通过学习率调度器实现高效快速收敛,优化帧移长度,并调节插值缩放因子以提升音质。同时,探索了泰米尔语和孟加拉语的潜在表示空间,用于声学单元发现与声纹转换任务。方法结合量化编码器与改进声码器,后者整合优化的帧移、调优插值因子及周期性学习率调度器,在英语、泰米尔语和孟加拉语数据集上均获得稳定优异表现。该方法能有效捕捉复杂语言特征,实现高质量语音重建,显著降低训练时长。
原文摘要 · Abstract (English)
This work addresses the persistent challenges of substantial training time and GPU resource requirements even when training lightweight encoder-vocoder models for Textless NLP. We reduce training steps significantly while improving performance by a) leveraging learning rate schedulers for efficient and faster convergence b) optimizing hop length and c) tuning the interpolation scale factors for better audio quality. Additionally, we explore the latent space representation for Indian languages such as Tamil and Bengali for the acoustic unit discovery and voice conversion task. Our approach leverages a quantized encoder architecture, in conjunction with a vocoder which utilizes the proposed mixture of optimized hop length, tuned interpolation scale factors and a cyclic learning rate scheduler. We obtain consistently good results across English, Tamil and Bengali datasets. The proposed method excels in capturing complex linguistic patterns, resulting in clear reconstructed audio during voice conversion with significantly reduced training time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。