针对印度语境下的噪声环境,优化神经网络以提升印地语语音分离与增强效果。
Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments
- 采用改进的DEMUCS模型,融合U-Net与LSTM结构提升语音清晰度。
- 在40万条印地语语音数据上训练,噪声环境下PESQ与STOI指标显著提升。
- 适配耳机等边缘设备,通过量化技术降低计算开销,适合实际部署。
本文针对噪声环境中印地语语音分离与增强的挑战,提出一种基于先进神经网络架构的优化方法,重点面向边缘设备部署。通过改进DEMUCS模型,引入U-Net与LSTM结构,利用包含40万条印地语语音片段的数据集进行训练,并结合ESC-50与MS-SNSD数据集进行声学环境多样化增强。在多种噪声条件下评估显示,该方法在PESQ与STOI指标上表现优异,尤其在极端噪声下优势明显。为适应资源受限设备如真无线耳机(TWS earbuds),研究进一步探索了量化技术以降低计算需求。本工作验证了定制化人工智能算法在印度语言语音处理中的有效性,并为边缘端架构优化提供未来方向。
原文摘要 · Abstract (English)
This paper addresses the challenges of Hindi speech separation and enhancement using advanced neural network architectures, with a focus on edge devices. We propose a refined approach leveraging the DEMUCS model to overcome limitations of traditional methods, achieving substantial improvements in speech clarity and intelligibility. The model is fine-tuned with U-Net and LSTM layers, trained on a dataset of 400,000 Hindi speech clips augmented with ESC-50 and MS-SNSD for diverse acoustic environments. Evaluation using PESQ and STOI metrics shows superior performance, particularly under extreme noise conditions. To ensure deployment on resource-constrained devices like TWS earbuds, we explore quantization techniques to reduce computational requirements. This research highlights the effectiveness of customized AI algorithms for speech processing in Indian contexts and suggests future directions for optimizing edge-based architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。