arXiv:2510.22241cs.SD2025-10被引 2

首个面向四通道全景声的低比特率神经编码器,保持方向信息准确。

FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss

  • 基于WavTokenizer扩展为四通道全景声,引入空间一致性损失。
  • 每秒仅75个离散码本,压缩至0.9 kbps,方向误差最低3.96°。
  • 适用于真实场景下的声音定位与检测任务,可直接用于下游应用。

神经音频编码器在单声道和立体声信号中已广泛应用,但空间音频仍鲜有研究。本文提出首个针对一阶全向声(FOA)的离散神经空间音频编码器。在WavTokenizer架构基础上,将其扩展以支持四通道FOA信号,并引入一种新颖的空间一致性损失,以在高度压缩的表示下保留重构信号的方向信息。该编码器将4通道、24 kHz的FOA音频压缩为每秒75个离散令牌,对应比特率为0.9 kbps。在模拟混响混合、无混响纯净语音以及使用真实房间冲激响应的FOA混合信号上的评估显示,重建精度良好,平均角度误差分别为13.76°、3.96°和25.83°。此外,由该编码器生成的离散潜在表示在下游空间音频任务中表现优异,已在STARSS23真实录音上验证了其在声音事件定位与检测中的有效性。

原文摘要 · Abstract (English)

Neural audio codecs have been widely studied for mono and stereo signals, but spatial audio remains largely unexplored. We present the first discrete neural spatial audio codec for first-order ambisonics (FOA). Building on the WavTokenizer architecture, we extend it to support four-channel FOA signals and introduce a novel spatial consistency loss to preserve directional cues in the reconstructed signals under a highly compressed representation. Our codec compresses 4-channel FOA audio at 24 kHz into 75 discrete tokens per second, corresponding to a bit rate of 0.9 kbps. Evaluations on simulated reverberant mixtures, non-reverberant clean speech, and FOA mixtures with real room impulse responses show accurate reconstruction, with mean angular errors of 13.76°, 3.96°, and 25.83°, respectively, across the three conditions. In addition, discrete latent representations derived from our codec provide useful features for downstream spatial audio tasks, as demonstrated on sound event localization and detection with STARSS23 real recordings.

空间音频神经编码全景声低比特率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。