arXiv:2410.11299cs.SDeess.AS2024-10被引 12

用扩散模型直接生成带空间感的音频,更真实更高效。

Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

  • 基于扩散变换器,端到端生成第一阶全向声场(FOA)。
  • 在两个数据集上优于传统方法,主观与客观指标均领先。
  • 适合做虚拟现实、游戏等需要精准空间音频的场景。

空间音频是打造沉浸式体验的关键。传统基于模拟的方法依赖专业知识,可扩展性差,且假设语义与空间信息独立。为此,我们探索端到端空间音频生成。提出新任务:根据声音类别和声源空间位置生成第一阶全向声场(FOA)。提出 Diff-SAGe,一种基于流的扩散-变压器模型,用于此任务。该模型采用复杂频谱图表示 FOA,保留对空间线索至关重要的相位信息。此外,多条件编码器将输入条件整合为统一表征,引导从噪声生成 FOA 波形。在两个数据集上的广泛评估表明,本方法在客观与主观评价上均持续优于传统模拟基线。

原文摘要 · Abstract (English)

Spatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalability, and assume independence between semantic and spatial information. To address these issues, we explore end-to-end spatial audio generation. We introduce and formulate a new task of generating first-order Ambisonics (FOA) given a sound category and sound source spatial location. We propose Diff-SAGe, an end-to-end, flow-based diffusion-transformer model for this task. Diff-SAGe utilizes a complex spectrogram representation for FOA, preserving the phase information crucial for accurate spatial cues. Additionally, a multi-conditional encoder integrates the input conditions into a unified representation, guiding the generation of FOA waveforms from noise. Through extensive evaluations on two datasets, we demonstrate that our method consistently outperforms traditional simulation-based baselines across both objective and subjective metrics.

空间音频扩散模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。