HyBeam融合麦克风与波束成形信号,提升可穿戴设备语音增强效果
HyBeam: Hybrid Microphone-Beamforming Array-Agnostic Speech Enhancement for Wearables
- 低频用原始麦克风信号,高频用波束成形信号,互补增强
- 在多种房间和佩戴配置下,性能优于单一方法
- 适合移动嵌入式设备,对麦克风阵列布局不敏感
语音增强是信号处理中的基础挑战,尤其在不同声学环境和麦克风配置下需保持鲁棒性。深度学习方法虽有效,但常假设固定阵列几何结构,限制了其在移动、嵌入式及可穿戴设备中的应用。现有无阵列依赖方法通常依赖原始麦克风信号或波束成形输出,但在阵列变化时均存在局限。我们提出HyBeam,一种混合框架:低频使用原始麦克风信号,高频使用波束成形信号,结合两者优势并保持高度阵列无关性。在多种房间及可穿戴阵列配置下的仿真表明,HyBeam在PESQ、STOI和SI-SDR指标上持续优于仅用麦克风或仅用波束成形的基线方法。频带分析显示,该方法在高频利用波束成形的方向性,在低频依赖麦克风线索,全频段表现均优于单一方法。
原文摘要 · Abstract (English)
Speech enhancement is a fundamental challenge in signal processing, particularly when robustness is required across diverse acoustic conditions and microphone setups. Deep learning methods have been successful for speech enhancement, but often assume fixed array geometries, limiting their use in mobile, embedded, and wearable devices. Existing array-agnostic approaches typically rely on either raw microphone signals or beamformer outputs, but both have drawbacks under changing geometries. We introduce HyBeam, a hybrid framework that uses raw microphone signals at low frequencies and beamformer signals at higher frequencies, exploiting their complementary strengths while remaining highly array-agnostic. Simulations across diverse rooms and wearable array configurations demonstrate that HyBeam consistently surpasses microphone-only and beamformer-only baselines in PESQ, STOI, and SI-SDR. A bandwise analysis shows that the hybrid approach leverages beamformer directivity at high frequencies and microphone cues at low frequencies, outperforming either method alone across all bands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。