arXiv:2510.00180eess.AScs.SD2025-10被引 1

用扩散模型从一阶音场生成三阶音场,提升沉浸感

DiffAU: Diffusion-Based Ambisonics Upscaling

  • 基于扩散模型设计级联音场升阶方法
  • 在无回声环境下多说话人场景下表现优异
  • 适合需要高质量空间音频的虚拟现实应用

空间音频通过还原三维声场增强沉浸感,其中球谐函数音场(Ambisonics)具备可扩展性。相较于高阶音场(HOA),一阶音场(FOA)在硬件采集与存储上更高效,但其空间分辨率低,限制了真实感,因此亟需音场升阶(AU)技术来提升音场阶数。本文提出 DiffAU,一种基于扩散模型的级联升阶方法,创新性地适配空间音频,可从一阶音场生成三阶音场。通过学习数据分布,DiffAU 在多种场景下快速且可靠地重建高阶音场。在多个说话人的无回声条件下实验表明,该方法在客观指标和主观感知上均表现良好。

原文摘要 · Abstract (English)

Spatial audio enhances immersion by reproducing 3D sound fields, with Ambisonics offering a scalable format for this purpose. While first-order Ambisonics (FOA) notably facilitates hardware-efficient acquisition and storage of sound fields as compared to high-order Ambisonics (HOA), its low spatial resolution limits realism, highlighting the need for Ambisonics upscaling (AU) as an approach for increasing the order of Ambisonics signals. In this work we propose DiffAU, a cascaded AU method that leverages recent developments in diffusion models combined with novel adaptation to spatial audio to generate 3rd order Ambisonics from FOA. By learning data distributions, DiffAU provides a principled approach that rapidly and reliably reproduces HOA in various settings. Experiments in anechoic conditions with multiple speakers, show strong objective and perceptual performance.

空间音频扩散模型音场升阶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。