arXiv:2504.04075eess.AScs.HC2025-04

实时模拟虚拟空间中第一人称语音回响,提升沉浸感。

Real-Time Auralization for First-Person Vocal Interaction in Immersive Virtual Environments

  • 用三维空间冲激响应实现实时语音音效渲染
  • 支持3自由度与5自由度的音视频同步交互
  • 适合虚拟音乐表演与语音研究场景

随着虚拟现实(VR)技术融合多模态反馈,真实空间的视听重现日益普遍。在VR体验中,用户语音是关键交互元素,如音乐演出和公开演讲应用。自我听觉感知对发声调节至关重要:歌唱或说话时,声音会受环境声学特性影响,从而调整发声参数。本技术报告提出一种实时声学化(auralization)流程,利用三维空间冲激响应(SIRs)支持需要第一人称语音交互的VR多模态研究应用。系统描述了冲激响应生成与渲染工作流、音视频融合方式,并解决延迟与计算开销问题。该系统允许用户在预设区域内从不同位置和朝向探索声学空间,支持3自由度(3DoF)与5自由度(5DoF)的音视频多模态感知,适用于科研与创意类VR应用。

原文摘要 · Abstract (English)

Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous applications rely on the user's voice as a key element of interaction, including music performances and public speaking applications. Self-perception of our voice plays a crucial role in vocal production. When singing or speaking, our voice interacts with the acoustic properties of the environment, shaping the adjustment of vocal parameters in response to the perceived characteristics of the space. This technical report presents a real-time auralization pipeline that leverages three-dimensional Spatial Impulse Responses (SIRs) for multimodal research applications in VR requiring first-person vocal interaction. It describes the impulse response creation and rendering workflow, the audio-visual integration, and addresses latency and computational considerations. The system enables users to explore acoustic spaces from various positions and orientations within a predefined area, supporting three and five Degrees of Freedom (3Dof and 5DoF) in audio-visual multimodal perception for both research and creative applications in VR.

虚拟现实语音交互声学仿真实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。