arXiv:2411.12008cs.SDcs.LG2024-11

用多通道RVQGAN压缩16声道三维音频,16kbps下仍保高音质。

Compression of Higher Order Ambisonics with Multichannel RVQGAN

  • 扩展RVQGAN支持16通道输入输出,不增加模型码率。
  • 在EigenScape数据集上,16kbps下听觉测试质量良好。
  • 适用于沉浸式音频编码,可迁移至其他多通道格式。

提出一种多通道延伸的RVQGAN神经编码方法,用于第三阶全向声场音频的数据驱动压缩。对生成器和判别器的输入输出层进行改造,无需增加模型码率即可处理16通道信号。设计了考虑沉浸式重放空间感知的损失函数,并采用单通道模型的迁移学习。在7.1.4沉浸式回放条件下,基于EigenScape数据库训练与测试的听觉实验表明,该方法可在16 kbps码率下有效编码场景式16通道全向声场内容,保持良好音质。模型具备拓展至其他内容类型和多通道格式的学习潜力。

原文摘要 · Abstract (English)

A multichannel extension to the RVQGAN neural coding method is proposed, and realized for data-driven compression of third-order Ambisonics audio. The input- and output layers of the generator and discriminator models are modified to accept multiple (16) channels without increasing the model bitrate. We also propose a loss function for accounting for spatial perception in immersive reproduction, and transfer learning from single-channel models. Listening test results with 7.1.4 immersive playback show that the proposed extension is suitable for coding scene-based, 16-channel Ambisonics content with good quality at 16 kbps when trained and tested on the EigenScape database. The model has potential applications for learning other types of content and multichannel formats.

音频压缩多通道全向声场神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。