arXiv:2511.11473cs.CLcs.SD2025-11EMNLP被引 2

主动识别对话对象,让助听器自动分离说话人。

Proactive Hearing Assistants that Isolate Egocentric Conversations

  • 用佩戴者自言作为锚点,结合对话轮换行为判断对方。
  • 在6.8小时真实场景数据上实现多对话场景下的精准分离。
  • 适合需要实时、无提示助听的听障人士或语音助手研发者。

我们提出主动式助听系统,无需用户明确指令即可自动识别并分离佩戴者正在对话的对象。系统基于佩戴者的双耳视角音频,利用其自身语音作为锚点,结合对话轮换行为与对话动态推断对话伙伴并抑制其他声音。为支持实时本地运行,设计双模型架构:轻量级流式模型每12.5毫秒运行一次,实现低延迟对话伙伴提取;较慢模型较少运行,捕捉更长时程对话动态。在来自11名参与者、总计6.8小时的真实世界双人与三人对话测试集上,系统展现出对多对话场景下对话伙伴的泛化识别与分离能力。本工作推动了能主动适应对话动态与参与度的助听设备发展。更多信息请见:https://proactivehearing.cs.washington.edu/

原文摘要 · Abstract (English)

We introduce proactive hearing assistants that automatically identify and separate the wearer's conversation partners, without requiring explicit prompts. Our system operates on egocentric binaural audio and uses the wearer's self-speech as an anchor, leveraging turn-taking behavior and dialogue dynamics to infer conversational partners and suppress others. To enable real-time, on-device operation, we propose a dual-model architecture: a lightweight streaming model runs every 12.5 ms for low-latency extraction of the conversation partners, while a slower model runs less frequently to capture longer-range conversational dynamics. Results on real-world 2- and 3-speaker conversation test sets, collected with binaural egocentric hardware from 11 participants totaling 6.8 hours, show generalization in identifying and isolating conversational partners in multi-conversation settings. Our work marks a step toward hearing assistants that adapt proactively to conversational dynamics and engagement. More information can be found on our website: https://proactivehearing.cs.washington.edu/

助听器语音分离主动交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。