提出新方法,高效模拟移动声源的动态混响,提升语音增强模型训练效果。
Fast Algorithm for Moving Sound Source
- 基于分层采样与时变延迟分解,构建物理合规的移动声场模型。
- 相比开源模型,幅度与相位还原更准确,实现实时动态数据生成。
- 适合语音增强、声学建模等需真实运动场景数据的研究者使用。
当前基于神经网络的语音处理系统普遍需要具备混响鲁棒性,因此训练需大量混响数据。现有方法多采用静态系统采样模拟动态系统,或依赖实录数据补充,但难以从根本上解决符合物理规律的运动数据模拟问题。针对移动场景下语音增强模型训练数据不足的核心难题,本文提出杨氏运动时空采样重构理论,实现对运动连续时变混响的高效模拟。该理论突破传统时变系统中静态图像源法(ISM)的局限,将移动像源的脉冲响应分解为线性时不变调制与离散时变分数延迟两部分,建立符合物理规律的移动声场模型。基于运动位移的带限特性,设计分层采样策略:低阶像源采用高采样率以保留细节,高阶像源采用低采样率以降低计算复杂度。进一步构建快速合成架构,实现实时模拟。实验表明,相较于开源模型,所提方法在移动场景下能更精确恢复幅度与相位变化,有效解决行业级运动声源数据模拟难题,为语音增强模型提供高质量动态训练数据。
原文摘要 · Abstract (English)
Modern neural network-based speech processing systems usually need to have reverberation resistance, so the training of such systems requires a large amount of reverberation data. In the process of system training, it is now more inclined to use sampling static systems to simulate dynamic systems, or to supplement data through actually recorded data. However, this cannot fundamentally solve the problem of simulating motion data that conforms to physical laws. Aiming at the core issue of insufficient training data for speech enhancement models in moving scenarios, this paper proposes Yang's motion spatio-temporal sampling reconstruction theory to realize efficient simulation of motion continuous time-varying reverberation. This theory breaks through the limitations of the traditional static Image-Source Method (ISM) in time-varying systems. By decomposing the impulse response of the moving image source into two parts: linear time-invariant modulation and discrete time-varying fractional delay, a moving sound field model conforming to physical laws is established. Based on the band-limited characteristics of motion displacement, a hierarchical sampling strategy is proposed: high sampling rate is used for low-order images to retain details, and low sampling rate is used for high-order images to reduce computational complexity. A fast synthesis architecture is designed to realize real-time simulation. Experiments show that compared with the open-source models, the proposed theory can more accurately restore the amplitude and phase changes in moving scenarios, solving the industry problem of motion sound source data simulation, and providing high-quality dynamic training data for speech enhancement models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。