arXiv:2601.12950eess.AS2026-01

将立体声直接生成7.1.4沉浸式音频,突破传统格式限制

ImmersiveFlow: Stereo-to-7.1.4 spatial audio generation with flow matching

  • 基于流匹配学习立体声到多通道声场的映射路径
  • 在预训练VAE隐空间中生成7.1.4音频,实现端到端合成
  • 显著提升声音外化感与听觉沉浸感,适合虚拟现实场景

沉浸式空间音频在增强现实、虚拟现实、家庭娱乐和汽车音响系统中日益重要。然而现有生成方法仍受限于低维格式,如双耳音频和一阶全向声学(FOA)。双耳渲染仅适用于耳机播放,而FOA存在空间混叠问题且高频分辨率不足。为克服这些局限,我们提出ImmersiveFlow,首个从立体声输入直接生成离散7.1.4格式空间音频的端到端生成框架。ImmersiveFlow利用流匹配技术,在预训练变分自编码器(VAE)隐空间中学习从立体声输入到多通道声场特征的映射轨迹。推理时,流匹配模型预测的隐空间特征经由VAE解码并转换为最终的7.1.4波形。全面的客观与主观评估表明,该方法生成的声音场更具感知丰富性,外化效果显著优于传统升频技术。代码与音频样例见:https://github.com/violet-audio/ImmersiveFlow。

原文摘要 · Abstract (English)

Immersive spatial audio has become increasingly critical for applications ranging from AR/VR to home entertainment and automotive sound systems. However, existing generative methods remain constrained to low-dimensional formats such as binaural audio and First-Order Ambisonics (FOA). Binaural rendering is inherently limited to headphone playback, while FOA suffers from spatial aliasing and insufficient resolution for high-frequency. To overcome these limitations, we introduce ImmersiveFlow, the first end-to-end generative framework that directly synthesizes discrete 7.1.4 format spatial audio from stereo input. ImmersiveFlow leverages Flow Matching to learn trajectories from stereo inputs to multichannel spatial features within a pretrained VAE latent space. At inference, the Flow Matching model predicted latent features are decoded by the VAE and converted into the final 7.1.4 waveform. Comprehensive objective and subjective evaluations demonstrate that our method produces perceptually rich sound fields and enhanced externalization, significantly outperforming traditional upmixing techniques. Code implementations and audio samples are provided at: https://github.com/violet-audio/ImmersiveFlow.

空间音频流匹配7.1.4音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。