用全景音频盲估声学参数,提升沉浸感。
Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
- 提出频带相关空间协方差特征,融合时频空信息。
- 误差减半,对10个频段的混响时间等参数估计更准。
- 适合做空间音频增强与虚拟现实音效开发的人看。
准确估计随频率变化的声学参数对实现真实感空间音频至关重要。本文提出一种统一框架,仅使用一阶全景音频(FOA)语音录音,盲估10个频段的混响时间(T60)、直达比(DRR)和清晰度(C50)。该框架引入新型特征——时频空协方差向量(SSCV),高效表征FOA信号的时间、频谱及空间特性。模型显著优于仅依赖频谱信息的单通道方法,三项参数估计误差均降低超过一半。此外,提出FOA-Conv3D网络,通过3D卷积编码器有效利用SSCV特征,性能优于传统CNN与循环卷积神经网络,各参数估计误差更低,解释方差比例(PoV)更高。
原文摘要 · Abstract (English)
Estimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60), direct-to-reverberant ratio (DRR), and clarity (C50) across 10 frequency bands using first-order Ambisonics (FOA) speech recordings as inputs. The proposed framework utilizes a novel feature named Spectro-Spatial Covariance Vector (SSCV), efficiently representing temporal, spectral as well as spatial information of the FOA signal. Our models significantly outperform existing single-channel methods with only spectral information, reducing estimation errors by more than half for all three acoustic parameters. Additionally, we introduce FOA-Conv3D, a novel back-end network for effectively utilising the SSCV feature with a 3D convolutional encoder. FOA-Conv3D outperforms the convolutional neural network (CNN) and recurrent convolutional neural network (CRNN) backends, achieving lower estimation errors and accounting for a higher proportion of variance (PoV) for all 3 acoustic parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。