无需训练数据,直接将单声道语音转为立体声,还能跨环境通用。
Zero-Shot Mono-to-Binaural Speech Synthesis
- 基于位置信息的几何时移与幅度缩放,无需参数即可生成初始立体声。
- 在标准数据集上感知效果媲美有监督方法,在新场景下更优。
- 适合需要快速部署、跨环境适配的音频增强与虚拟现实应用。
我们提出 ZeroBAS,一种无需任何立体声训练数据的神经方法,可从单声道录音和位置信息中合成立体声音频。据我们所知,这是首个公开发布的零样本单声道到立体声音频合成方法。具体而言,我们发现基于声源位置的无参数几何时移与幅度缩放即可生成初始立体声,再通过迭代应用预训练去噪声声码器进行优化。此外,该方法在不同房间条件下表现出良好泛化能力,我们为此引入新数据集 TUT Mono-to-Binaural,用于评估现有方法在未见条件下的性能。我们的零样本方法在标准数据集上的感知质量与监督方法相当,甚至在分布外的 TUT Mono-to-Binaural 数据集上表现更优。结果表明,预训练生成模型与零样本学习在实现鲁棒立体声合成方面具有巨大潜力。
原文摘要 · Abstract (English)
We present ZeroBAS, a neural method to synthesize binaural audio from monaural audio recordings and positional information without training on any binaural data. To our knowledge, this is the first published zero-shot neural approach to mono-to-binaural audio synthesis. Specifically, we show that a parameter-free geometric time warping and amplitude scaling based on source location suffices to get an initial binaural synthesis that can be refined by iteratively applying a pretrained denoising vocoder. Furthermore, we find this leads to generalization across room conditions, which we measure by introducing a new dataset, TUT Mono-to-Binaural, to evaluate state-of-the-art monaural-to-binaural synthesis methods on unseen conditions. Our zero-shot method is perceptually on-par with the performance of supervised methods on the standard mono-to-binaural dataset, and even surpasses them on our out-of-distribution TUT Mono-to-Binaural dataset. Our results highlight the potential of pretrained generative audio models and zero-shot learning to unlock robust binaural audio synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。