arXiv:2609.08542eess.AS2026-09

通过估计相对混响响应,实现高效沉浸式音频编码

Spatial Audio Coding Through Relative Room Impulse Response Estimation

论文配图:Spatial Audio Coding Through Relative Room Impulse Response Estimation
图 1 · 摘自论文原文
  • 基于波束成形的HOA信号盲估相对混响响应
  • 单传输通道下压缩率优于IVAS,音质相当或更优
  • 适合对房间声学建模要求高的沉浸式音频应用

沉浸式虚拟听觉依赖高阶全向声(HOA)等空间音频技术,将声音场景表示为多通道信号。随着空间分辨率提升,通道数增加,带宽受限网络下高效压缩成为关键。理想情况下,沉浸式音频编码目标码率应接近当前VoLTE语音服务分配的25 kbps。现有参数化编解码器如新标准的沉浸式语音与音频服务(IVAS)通过传输空间元数据和少量传输通道实现压缩。但研究表明,IVAS在混响内容下性能下降,尤其在低码率时,表明其难以准确建模房间声学。本文提出一种新型HOA编码方案,基于相对空间混响响应(ReSRIR)的显式与盲估计,以波束成形的HOA信号作为参考信号。利用估计出的ReSRIR结构与稀疏性,推导出高效的参数化表示用于沉浸式音频编码。实验表明,在单传输通道场景下,所提方法压缩率高于IVAS,同时保持相当或略优的音质。

原文摘要 · Abstract (English)

Immersive virtual listening relies on spatial audio technologies such as Higher-Order Ambisonics (HOA), which represent sound scenes as multichannel signals. As the desired spatial resolution increases, so does the number of channels, making efficient compression essential for transmission over bandwidth-limited networks. Moreover, to facilitate deployment by network operators, the target bitrate for immersive audio coding should ideally remain close to the 25 kbps currently allocated to VoLTE audio services. State-of-the-art parametric codecs, such as the recently standardized Immersive Voice and Audio Services (IVAS) codec, achieve compression by transmitting spatial metadata together with a reduced number of transport channels. However, recent studies have shown that IVAS performance degrades on reverberant content, particularly at low bitrates, a limitation that suggests its inability to accurately model room acoustics. In this paper, we propose a novel HOA coding scheme based on the explicit and blind estimation of the Relative Spatial Room Impulse Response (ReSRIR), using a beamformed version of the HOA signal as a reference signal. By exploiting the structure and sparsity of the estimated ReSRIR, we derive an efficient parametric representation for immersive audio coding. Experimental evaluations show that the proposed method achieves higher compression than IVAS in the single-transport-channel regime, while maintaining comparable to slightly better quality.

空间音频混响建模参数编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。