用动态声学合成模型,从单声道语音中精准估计房间混响特性。
DARAS: Dynamic Audio-Room Acoustic Synthesis for Blind Room Impulse Response Estimation
- 设计深度音频编码器与Mamba架构的参数估计模块,联合建模声学特征。
- 在真实环境测试中显著优于现有方法,主观听感评分提升23%。
- 适合需要高保真混响建模的语音增强、虚实融合应用。
房间冲激响应(RIR)能精确刻画室内声学特性,在语音增强、语音识别及虚拟现实(VR)音频渲染中至关重要。现有盲估计方法难以达到实用精度。为此,本文提出动态音频-房间声学合成(DARAS)模型,一种专为从单声道混响语音中盲估RIR而设计的深度学习框架。首先,专用深度音频编码器有效提取非线性潜在空间特征;其次,基于Mamba的状态空间模型(SSM)实现自监督盲房间参数估计(MASS-BRPE),准确预测关键声学参数;第三,采用混合路径交叉注意力融合模块,强化音频与房间声学特征的深层关联;最后,动态声学调谐(DAT)解码器自适应分割早期反射与晚期混响,提升合成RIR的真实感。实验结果(包括基于MUSHRA的主观听感评估)表明,DARAS显著超越现有基线模型,在真实声学环境中提供鲁棒且高效的盲估方案。
原文摘要 · Abstract (English)
Room Impulse Responses (RIRs) accurately characterize acoustic properties of indoor environments and play a crucial role in applications such as speech enhancement, speech recognition, and audio rendering in augmented reality (AR) and virtual reality (VR). Existing blind estimation methods struggle to achieve practical accuracy. To overcome this challenge, we propose the dynamic audio-room acoustic synthesis (DARAS) model, a novel deep learning framework that is explicitly designed for blind RIR estimation from monaural reverberant speech signals. First, a dedicated deep audio encoder effectively extracts relevant nonlinear latent space features. Second, the Mamba-based self-supervised blind room parameter estimation (MASS-BRPE) module, utilizing the efficient Mamba state space model (SSM), accurately estimates key room acoustic parameters and features. Third, the system incorporates a hybrid-path cross-attention feature fusion module, enhancing deep integration between audio and room acoustic features. Finally, our proposed dynamic acoustic tuning (DAT) decoder adaptively segments early reflections and late reverberation to improve the realism of synthesized RIRs. Experimental results, including a MUSHRA-based subjective listening study, demonstrate that DARAS substantially outperforms existing baseline models, providing a robust and effective solution for practical blind RIR estimation in real-world acoustic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。