arXiv:2509.01399cs.SDcs.AI2025-09中稿 · Interspeech 2025被引 1

轻量级语音分离模型,提升车内多说话人识别准确率

CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays

  • 用空间特征增强掩码估计,提升语音与噪声区分能力
  • 推理时采用MVDR算法,降低语音失真,更适配语音识别
  • 结合仿真与真实混响数据增强,改善边界区域定位效果

从多个说话人中分离重叠语音对人车交互至关重要。本文提出CabinSep,一种轻量级神经掩码式最小方差无失真响应(MVDR)语音分离方法,旨在降低后端自动语音识别(ASR)模型的错误率。贡献有三:首先,利用通道信息提取空间特征,提升语音与噪声掩码估计精度;其次,在推理阶段使用MVDR,减少语音失真,使其更符合ASR需求;第三,引入结合仿真与真实混响响应(IRs)的数据增强方法,改善声源定位在区域边界的表现,进一步降低识别错误。模型计算复杂度仅0.4 GMACs,相较于当前最优的DualSep模型,在真实录音数据集上实现17.5%的相对识别错误率下降。演示视频见:https://cabinsep.github.io/cabinsep/

原文摘要 · Abstract (English)

Separating overlapping speech from multiple speakers is crucial for effective human-vehicle interaction. This paper proposes CabinSep, a lightweight neural mask-based minimum variance distortionless response (MVDR) speech separation approach, to reduce speech recognition errors in back-end automatic speech recognition (ASR) models. Our contributions are threefold: First, we utilize channel information to extract spatial features, which improves the estimation of speech and noise masks. Second, we employ MVDR during inference, reducing speech distortion to make it more ASR-friendly. Third, we introduce a data augmentation method combining simulated and real-recorded impulse responses (IRs), improving speaker localization at zone boundaries and further reducing speech recognition errors. With a computational complexity of only 0.4 GMACs, CabinSep achieves a 17.5% relative reduction in speech recognition error rate in a real-recorded dataset compared to the state-of-the-art DualSep model. Demos are available at: https://cabinsep.github.io/cabinsep/.

语音分离车载系统MVDRASR优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。