用声学微结构让手机定向拾音,嘈杂环境也能清晰录音。
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures

- 通过生物启发的声学微结构,被动实现声音方向感知。
- 在30°方向上提升信号质量5.0 dB,性能超传统五麦克风阵列。
- 端到端神经网络实时运行,适配低成本有线耳机手机使用。
设想将手机放在嘈杂餐厅的桌上,仍能清晰捕捉周围朋友的对话,或在混响强烈的礼堂中清晰记录讲师发言。我们提出SonicSieve,首个基于生物启发声学微结构的智能手机智能定向语音提取系统。其无源设计通过物理结构嵌入方向性线索,无需额外电子元件,可直接安装于低成本有线耳机的内置麦克风上。系统采用端到端神经网络,在移动设备上实时处理原始音频混合信号。实验表明,当聚焦于30°角度区域时,信号质量提升5.0 dB;且仅使用两个麦克风的系统性能优于传统五麦克风阵列。
原文摘要 · Abstract (English)
Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30° angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。