一个能处理单声道到5.1环绕声的通用神经音频编解码器。
VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround Sound
- 统一编码解码结构,支持多通道实时切换。
- 在单声道、立体声和5.1环绕声下均保持高保真度。
- 适合需要跨通道兼容的语音与影视音频应用。
我们提出VCNAC,一种可变通道的神经音频编解码器。其采用单一编码器与解码器参数化设计,原生支持从单声道语音到影院级5.1环绕声的多通道推理。通过通道兼容性目标,确保多通道内容在降解至少通道时仍保持感知质量。共享表示使生成式语言模型仅需一套代码本即可训练,并在推理时实现模态与通道配置的可扩展性。基于客观空间音频指标和主观听觉测试评估,该统一方法在单声道、立体声及环绕声配置中均维持高重建质量。
原文摘要 · Abstract (English)
We present VCNAC, a variable channel neural audio codec. Our approach features a single encoder and decoder parametrization that enables native inference for different channel setups, from mono speech to cinematic 5.1 channel surround audio. Channel compatibility objectives ensure that multi-channel content maintains perceptual quality when decoded to fewer channels. The shared representation enables training of generative language models on a single set of codebooks while supporting inference-time scalability across modalities and channel configurations. Evaluation using objective spatial audio metrics and subjective listening tests demonstrates that our unified approach maintains high reconstruction quality across mono, stereo, and surround audio configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。