智能眼镜语音交互降噪新方案,有效抑制旁人说话干扰
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
- 多麦克风输入融合采用Tri-Mamba架构,支持实时处理
- 帧级和语句级联合优化,噪声环境下误识别率降低4.95%
- 专为智能眼镜设计,适合嘈杂场景下的语音交互应用
智能眼镜正成为接入大语言模型的下一代人机接口,但在真实噪声环境中实现可靠交互仍面临挑战,尤其受旁人说话干扰影响。本文提出一种新型多麦克风Whisper框架(MMW),包含三项关键创新:首先,基于Tri-Mamba架构设计的混合模块,在原始波形层面高效融合多通道音频,同时保持流式处理兼容性;其次,设计帧级语音分离的Mamba层,增强帧级旁人语音抑制能力,提升Whisper模型微调效率;第三,采用多尺度组相对策略优化(GRPO),联合优化帧级与语句级旁人语音抑制。实验表明,所提MMW系统在噪声条件下可将词错误率(WER)降低4.95%。
原文摘要 · Abstract (English)
Smart glasses are increasingly positioned as the next-generation interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in real-world noisy environments remains a major challenge, particularly due to interference from side speech. In this work, we introduce a novel side-talk rejection multi-microphone Whisper (MMW) framework for smart glasses, incorporating three key innovations. First, we propose a Mix Block based on a Tri-Mamba architecture to effectively fuse multi-channel audio at the raw waveform level, while maintaining compatibility with streaming processing. Second, we design a Frame Diarization Mamba Layer to enhance frame-level side-talk suppression, facilitating more efficient fine-tuning of Whisper models. Third, we employ a Multi-Scale Group Relative Policy Optimization (GRPO) strategy to jointly optimize frame-level and utterance-level side speech suppression. Experimental evaluations demonstrate that the proposed MMW system can reduce the word error rate (WER) by 4.95\% in noisy conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。