让神经网络学会通用麦克风阵列的声场编码,无需重训练
Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
- 用双路径编码器分别处理阵列几何与信号,几何信息引导信号处理
- 在无混响场景中全频段优于传统方法,有混响时频段依赖性提升
- 适合需要快速适配新麦克风阵列的实时空间音频应用
使用深度神经网络(DNN)将麦克风阵列(MA)信号编码为Ambisonics空间音频格式,可突破传统方法的局限,但现有DNN方法需针对每种MA单独训练。本文提出一种DNN-based Ambisonics编码方法,可泛化至训练中未见的任意MA几何结构。该方法输入包括MA几何与信号,采用多层级编码器,包含独立的几何与信号路径,几何特征在每一层级影响信号编码器。在模拟的无混响与混响条件下,单源和双源场景中验证了该方法的有效性。结果表明,在干爽场景中全频段性能优于传统编码;在混响场景中,改进具有频率依赖性。
原文摘要 · Abstract (English)
Using deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained separately for each MA. This paper proposes a DNN-based method for Ambisonics encoding that can generalize to arbitrary MA geometries unseen during training. The method takes as inputs the MA geometry and MA signals and uses a multi-level encoder consisting of separate paths for geometry and signal data, where geometry features inform the signal encoder at each level. The method is validated in simulated anechoic and reverberant conditions with one and two sources. The results indicate improvement over conventional encoding across the whole frequency range for dry scenes, while for reverberant scenes the improvement is frequency-dependent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。