动态声音模拟器可生成三维移动声源的多通道音频,适合训练定位算法。
DynamicSound simulator for simulating moving sources and microphone arrays
- 支持3D空间中连续移动的声源与任意麦克风阵列仿真
- 精确建模传播延迟、多普勒效应和环境反射,生成时序一致音频
- 开源可复现,适用于空间音频与声源定位算法开发
开发声音分类、检测与定位算法需要大量灵活且真实的音频数据,尤其在使用现代机器学习和波束成形技术时。然而,现有声学仿真工具多针对室内环境,仅支持静态声源,难以应对移动声源、移动麦克风或远距离传播场景。本文提出DynamicSound——一个开源声学仿真框架,可生成一个或多个在三维空间中连续移动的声源所对应的多通道音频,并由任意配置的麦克风阵列录制。该模型显式考虑有限传播延迟、多普勒效应、距离衰减、空气吸收及平面表面的一阶反射,生成时间上一致的空间音频信号。与传统单声道或立体声仿真器不同,该系统可合成任意数量虚拟麦克风的音频,准确还原麦克风间的时间差、电平差及环境引起的频谱着色。与现有开源工具对比评估表明,生成信号在不同声源位置和声学条件下均保持高空间保真度。该开源框架通过在受控且可重复条件下生成真实感多通道音频,为现代空间音频与声源定位算法的开发、训练与评估提供了灵活可靠的工具。
原文摘要 · Abstract (English)
Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing acoustic simulators are tailored for indoor environments and are limited to static sound sources, making them unsuitable for scenarios involving moving sources, moving microphones, or long-distance propagation. This paper presents DynamicSound an open-source acoustic simulation framework for generating multichannel audio from one or more sound sources with the possibility to move them continuously in three-dimensional space and recorded by arbitrarily configured microphone arrays. The proposed model explicitly accounts for finite sound propagation delays, Doppler effects, distance-dependent attenuation, air absorption, and first-order reflections from planar surfaces, yielding temporally consistent spatial audio signals. Unlike conventional mono or stereo simulators, the proposed system synthesizes audio for an arbitrary number of virtual microphones, accurately reproducing inter-microphone time delays, level differences, and spectral coloration induced by the environment. Comparative evaluations with existing open-source tools demonstrate that the generated signals preserve high spatial fidelity across varying source positions and acoustic conditions. By enabling the generation of realistic multichannel audio under controlled and repeatable conditions, the proposed open framework provides a flexible and reproducible tool for the development, training, and evaluation of modern spatial audio and sound-source localization algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。