arXiv:2608.25404eess.AS2026-08

实时生成空间音频的因果模型,兼顾音质与空间感。

CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation

论文配图:CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
图 1 · 摘自论文原文
  • 用动态声源-听者位置融合多通道频谱,提升空间感知
  • 引入空间一致性判别器,优化通道间线索,空间保真度↑12%
  • 严格因果设计支持低延迟流式推理,内存开销恒定

空间音频声码器可将生成模型输出的梅尔频谱转换为包含空间信息的音频波形。现有神经声码器多针对单声道设计,直接扩展至空间音频会因忽略通道间线索而降低空间质量。本文提出CSAVocoder,一种基于因果GAN的空间音频声码器,联合优化波形保真度与空间渲染效果。其框架引入空间适配器(Spatial Adaptor),融合多通道梅尔频谱与动态声源-听者姿态信息,并设计空间一致性判别器,监督通道间线索。为满足实时性要求,采用严格因果、状态化的生成器,支持高效流式推理且内存开销恒定。在大规模空间音频数据集上的实验表明,CSAVocoder在保持竞争性音频质量的同时,显著提升空间保真度,具备实时性能。

原文摘要 · Abstract (English)

Spatial audio vocoders are able to convert mel-spectrograms produced by generative models into spatial audio waveforms. Most neural vocoders are designed for monaural audio, and direct extensions to spatial audio can degrade spatial quality by ignoring inter-channel cues. We present CSAVocoder, a causal GAN-based spatial audio vocoder that jointly optimizes waveform fidelity and spatial rendering. Our framework introduces a Spatial Adaptor that fuses multi-channel mel-spectrograms with dynamic source-listener pose information, together with a spatial consistency discriminator that supervises inter-channel cues. To meet real-time requirements, we design a strictly causal, stateful generator that supports efficient streaming inference with constant memory overhead. Experiments on large-scale spatial audio datasets show that CSAVocoder improves spatial fidelity at competitive audio quality and real-time performance.

空间音频声码器生成模型实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。