arXiv:2504.18715cs.CLcs.SD2025-04中稿 · CHI2025被引 11

让耳机在嘈杂环境中实时翻译语音并保持声音方向感。

Spatial Speech Translation: Translating Across Space With Binaural Hearables

  • 用多技术协同实现空间语音翻译,保留说话人方向和音色。
  • 在强干扰下仍达22.01的BLEU得分,支持实时运行。
  • 适合需要真实环境语音翻译的可穿戴设备用户。

设想身处一个多人讲不同语言的嘈杂环境中,佩戴耳戴设备能将周围语音转化为你的母语,同时保留每位说话人的空间位置信息。我们提出空间语音翻译这一新概念,使耳戴设备可在翻译环境中的语音时,维持各说话人的方向与独特音色特征,并在双耳输出中体现。为实现该目标,我们解决了盲源分离、定位、实时表达性翻译及双耳渲染等一系列技术挑战,确保翻译后语音的空间方向性,并在Apple M2芯片上实现实时推理。原型头戴设备的验证表明,相比现有模型在干扰环境下失效的情况,我们的系统在存在强烈背景干扰时仍能达到最高22.01的BLEU得分。用户研究进一步证实,系统能在未见过的真实混响环境中有效进行空间化语音渲染。这项工作标志着首次将空间感知融入语音翻译的关键一步。

原文摘要 · Abstract (English)

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech translation, a novel concept for hearables that translate speakers in the wearer's environment, while maintaining the direction and unique voice characteristics of each speaker in the binaural output. To achieve this, we tackle several technical challenges spanning blind source separation, localization, real-time expressive translation, and binaural rendering to preserve the speaker directions in the translated audio, while achieving real-time inference on the Apple M2 silicon. Our proof-of-concept evaluation with a prototype binaural headset shows that, unlike existing models, which fail in the presence of interference, we achieve a BLEU score of up to 22.01 when translating between languages, despite strong interference from other speakers in the environment. User studies further confirm the system's effectiveness in spatially rendering the translated speech in previously unseen real-world reverberant environments. Taking a step back, this work marks the first step towards integrating spatial perception into speech translation.

语音翻译空间音频可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。