动态追踪移动说话者,实时增强特定方向声音。
Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers
- 用隐式定位在线组合多个双耳滤波器。
- 支持实时跟踪移动声源,提升语音聚焦与降噪效果。
- 适用于虚拟现实中的世界锁定音频,兼容多种麦克风阵列。
我们提出一种新型的专家混合框架,用于双耳信号匹配中的视场增强。该方法实现动态空间音频渲染,可适应连续说话者运动,让用户强调或抑制特定方向的声音,同时保持自然的双耳线索。与依赖显式到达方向估计或在阿姆比森斯域操作的传统方法不同,我们的信号依赖型框架通过隐式定位在线组合多个双耳滤波器,实现对移动声源的实时跟踪与增强,适用于语音聚焦、降噪及增强现实和虚拟现实中的世界锁定音频。该方法不依赖阵列几何结构,为下一代消费级音频设备提供灵活的空间音频采集与个性化回放方案。
原文摘要 · Abstract (English)
We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。