IVAS用参数编码在低码率下高效还原多声道音频空间感
Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates
- 通过方向信息、主对象索引和功率比生成参数侧信息
- 24.4~32 kbit/s可还原3~4个任意位置音频对象的空间声场
- 适合对低码率沉浸式音频有需求的实时通信场景
最新标准化的3GPP沉浸式语音与音频服务(IVAS)编码器包含一种用于在低码率下高效编码多个音频对象的参数模式。该模式从对象元数据和输入音频对象中获取参数侧信息,包括方向信息、两个主对象的索引以及这两个主对象间的功率比。这些侧信息连同立体声混合信号一起传输至解码器。在IVAS中,参数化对象编码可在24.4或32 kbit/s的码率下,实现三个或四个任意放置音频对象的传输,并忠实重建原始音频场景的空间图像。主观听音测试表明,相比使用增强语音服务(EVS)独立编码音频对象,IVAS在更低码率和更低成本下提供了相当的沉浸式体验。
原文摘要 · Abstract (English)
The recently standardized 3GPP codec for Immersive Voice and Audio Services (IVAS) includes a parametric mode for efficiently coding multiple audio objects at low bit rates. In this mode, parametric side information is obtained from both the object metadata and the input audio objects. The side information comprises directional information, indices of two dominant objects, and the power ratio between these two dominant objects. It is transmitted to the decoder along with a stereo downmix. In IVAS, parametric object coding allows for transmitting three or four arbitrarily placed objects at bit rates of 24.4 or 32 kbit/s and faithfully reconstructing the spatial image of the original audio scene. Subjective listening tests confirm that IVAS provides a comparable immersive experience at lower bit rate and complexity compared to coding the audio objects independently using Enhanced Voice Services (EVS).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。