arXiv:2502.08848cs.HCcs.SD2025-02中稿 · CHI 2025被引 3

让手机语音转文字识别说话人方向,提升多人对话可读性

SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization

  • 用四个麦克风实时定位声音方向,生成箭头等视觉引导
  • 用户调研显示94%认为方向提示对群聊有明显帮助
  • 适合听障人士、会议记录者和多语言使用者

移动端语音转文字功能在听力辅助、语言翻译、笔记记录和会议纪要中已显成效。但我们的大规模调查(n=263)发现,在多人对话中无法区分说话人及其方向,严重影响使用体验。SpeechCompass通过实时多麦克风语音定位技术,实现说话方向的可视化分离与引导(如箭头指示)。我们设计了高效低功耗的音频定位算法,并部署于集成四麦克风的定制感知硬件上,运行在低功耗微控制器上,经技术评估验证其性能。基于更大规模调研(n=494),我们开展了面对面群组对话研究,邀请八位频繁使用移动语音转文字的用户测试五种可视化风格。结果表明,所有参与者一致认可去重和方向提示的价值,认为定向引导显著提升多人对话场景下的可读性和可用性。

原文摘要 · Abstract (English)

Speech-to-text capabilities on mobile devices have proven helpful for hearing and speech accessibility, language translation, note-taking, and meeting transcripts. However, our foundational large-scale survey (n=263) shows that the inability to distinguish and indicate speaker direction makes them challenging in group conversations. SpeechCompass addresses this limitation through real-time, multi-microphone speech localization, where the direction of speech allows visual separation and guidance (e.g., arrows) in the user interface. We introduce efficient real-time audio localization algorithms and custom sound perception hardware running on a low-power microcontroller and four integrated microphones, which we characterize in technical evaluations. Informed by a large-scale survey (n=494), we conducted an in-person study of group conversations with eight frequent users of mobile speech-to-text, who provided feedback on five visualization styles. The value of diarization and visualizing localization was consistent across participants, with everyone agreeing on the value and potential of directional guidance for group conversations.

语音识别多麦克风方向定位无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。