用结构感知约束提升声学环境的神经压缩效率
Deep Neural Compression for RIR-Characterized Acoustic Environments with Structure-Aware Constraints

- 在全局衰减与局部能量上施加结构约束,增强压缩保真度
- 375 bps下重建误差更低,混响语音感知一致性更优
- 适合沉浸式音频、声学建模等需高保真混响数据场景
房间脉冲响应(RIR)通过捕捉声音在封闭空间内的传播与衰减特性,表征房间声学环境。在沉浸式音频渲染等应用中,精确声学重建通常依赖于空间密集采样的RIR数据,导致数据量庞大,存储负担重。尽管近期神经音频编解码器为低比特率压缩提供了有效框架,但其训练目标主要针对语音和通用音频,难以匹配RIR的声学特性。为此,我们提出一种基于EnCodec的神经RIR压缩方法,从两个层面引入结构感知约束:在RIR层面,通过能量衰减曲线(EDC)正则化和短时窗能量约束,分别控制全局衰减行为与局部能量分布;在混响语音层面,引入混响语音监督,确保重构RIR生成的混响语音具有一致性。实验表明,在375 bps低比特率下,该方法相比面向音频的编解码器,实现了更低的RIR重建误差与更优的混响语音感知一致性。
原文摘要 · Abstract (English)
Room impulse responses (RIRs) characterize the acoustic environment of a room by capturing how sound propagates and decays within an enclosed space. In applications such as immersive audio rendering, accurate acoustic reconstruction often relies on spatially densely sampled RIRs. This consequently gives rise to a large volume of RIR data, imposing a substantial burden on storage. Although recent neural audio codecs provide an effective framework for low-bitrate compression, their training objectives are mainly tailored to speech and general audio, and are therefore not well aligned with the acoustic characteristics of RIRs. Therefore, we propose an EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels. Specifically, at the RIR level, structure-aware constraints are imposed on the global decay behavior and local energy distribution of RIRs through energy decay curve (EDC) regularization and a short-time window energy constraint, while at the reverberant-speech level, reverberant-speech supervision is further introduced to constrain the consistency of the reverberant speech generated by the reconstructed RIRs. Experimental results show that, at a low bitrate of 375 bps, the proposed method achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。