轻量级语音分离架构,实现实时车内多区语音分离
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
- 融合梅尔谱图与双耳相位差,压缩空间信息提取开销
- 仅0.56G MACs、RTF 0.37,复杂场景下仍保持高精度
- 适合车载实时语音交互,部署门槛低
车内多区语音分离技术对人车交互至关重要。尽管先前的SpatialNet已取得显著效果,但其高计算成本限制了车辆中的实时应用。为此,本文提出LSZone,一种面向实时车内多区语音分离的轻量级空间信息建模架构。设计了融合梅尔谱图与双耳相位差(IPD)的空间信息提取-压缩模块(SpaIEC),在保持性能的同时降低计算负担。此外,引入极轻量级的Conv-GRU跨频带窄带处理模块(CNP),高效建模空间信息。实验表明,LSZone在复杂噪声和多人说话场景下表现优异,仅需0.56G MACs、实时因子(RTF)为0.37,具备实际部署潜力。
原文摘要 · Abstract (English)
In-car multi-zone speech separation, which captures voices from different speech zones, plays a crucial role in human-vehicle interaction. Although previous SpatialNet has achieved notable results, its high computational cost still hinders real-time applications in vehicles. To this end, this paper proposes LSZone, a lightweight spatial information modeling architecture for real-time in-car multi-zone speech separation. We design a spatial information extraction-compression (SpaIEC) module that combines Mel spectrogram and Interaural Phase Difference (IPD) to reduce computational burden while maintaining performance. Additionally, to efficiently model spatial information, we introduce an extremely lightweight Conv-GRU crossband-narrowband processing (CNP) module. Experimental results demonstrate that LSZone, with a complexity of 0.56G MACs and a real-time factor (RTF) of 0.37, delivers impressive performance in complex noise and multi-speaker scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。