arXiv:2511.09525cs.HCcs.SD2025-11被引 1

空间音频提升实时翻译理解力,让跨国会议更清晰自然。

Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings

  • 用空间音频模拟说话人位置,增强翻译语音的方位感。
  • 空间音效使理解准确率提升至非空间音的两倍(68%→136%)。
  • 适合远程会议系统、多语言协作平台的设计优化。

虚拟会议中的语言障碍仍是全球协作的主要挑战。实时翻译虽具潜力,但现有方案常忽略感知线索。本研究探究空间音频渲染对翻译语音在多语言会议中理解度、认知负荷与用户体验的影响。实验采用8名双语同谋者和47名参与者,模拟包含希腊语、卡纳达语、中文和乌克兰语的全球团队会议,这些语言在语法、文字系统及资源丰富度上差异显著。参与者体验四种音频条件:带/不带混响的空间音频,以及两种非空间配置(双耳、单声道)。通过测量理解准确率、工作量评分、满意度与质性反馈发现,空间渲染翻译的理解准确率较非空间音频翻倍;当空间线索与语音音色区分同时存在时,用户感知更清晰、参与感更强。研究为实时翻译在会议平台中的集成提供设计启示,推动沉浸式跨语言通信的发展。

原文摘要 · Abstract (English)

Language barriers in virtual meetings remain a persistent challenge to global collaboration. Real-time translation offers promise, yet current integrations often neglect perceptual cues. This study investigates how spatial audio rendering of translated speech influences comprehension, cognitive load, and user experience in multilingual meetings. We conducted a within-subjects experiment with 8 bilingual confederates and 47 participants simulating global team meetings with English translations of Greek, Kannada, Mandarin Chinese, and Ukrainian - languages selected for their diversity in grammar, script, and resource availability. Participants experienced four audio conditions: spatial audio with and without background reverberation, and two non-spatial configurations (diotic, monaural). We measured listener comprehension accuracy, workload ratings, satisfaction scores, and qualitative feedback. Spatially-rendered translations doubled comprehension compared to non-spatial audio. Participants reported greater clarity and engagement when spatial cues and voice timbre differentiation were present. We discuss design implications for integrating real-time translation into meeting platforms, advancing inclusive, cross-language communication in telepresence systems.

空间音频实时翻译虚拟会议多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。