arXiv:2505.21004cs.SD2025-05被引 1

多耳机会话增强系统,让嘈杂环境中对话更清晰。

CoHear: Conversation Enhancement via Multi-Earphone Collaboration

  • 通过协作网络协议实现耳戴设备实时协同
  • 语音质量提升最高达8.8 dB,群组识别准确率超90%
  • 适合嘈杂环境下的会议、社交场景使用

在会议等嘈杂场所,背景噪音、多人重叠说话使对话难以听清,加剧了‘鸡尾酒会聋’现象。我们提出ClearSphere系统,通过多耳机协作实现对话层面的语音增强。该系统需对对话成员进行整体建模,并有效从混合声音中提取目标语音。ClearSphere通过两项关键贡献实现:1)面向对话的网络协议;2)鲁棒的目标对话提取模型。网络协议支持移动、无基础设施的耳机设备协同;语音提取模型能以高效带宽方式利用中继音频。在真实实验与仿真中评估显示,该系统群组形成准确率超过90%,语音质量相比最先进基线最高提升8.8 dB,且可在移动端实现实时运行。20人用户研究显示,ClearSphere得分显著高于基线,具备良好可用性。

原文摘要 · Abstract (English)

In crowded places such as conferences, background noise, overlapping voices, and lively interactions make it difficult to have clear conversations. This situation often worsens the phenomenon known as "cocktail party deafness." We present ClearSphere, the collaborative system that enhances speech at the conversation level with multi-earphones. Real-time conversation enhancement requires a holistic modeling of all the members in the conversation, and an effective way to extract the speech from the mixture. ClearSphere bridges the acoustic sensor system and state-of-the-art deep learning for target speech extraction by making two key contributions: 1) a conversation-driven network protocol, and 2) a robust target conversation extraction model. Our networking protocol enables mobile, infrastructure-free coordination among earphone devices. Our conversation extraction model can leverage the relay audio in a bandwidth-efficient way. ClearSphere is evaluated in both real-world experiments and simulations. Results show that our conversation network obtains more than 90\% accuracy in group formation, improves the speech quality by up to 8.8 dB over state-of-the-art baselines, and demonstrates real-time performance on a mobile device. In a user study with 20 participants, ClearSphere has a much higher score than baseline with good usability.

语音增强多设备协同耳戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。