提出新编码方法,同时实现语音定位与分离,性能优于单一任务优化方案。
Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask Estimation
- 用掩码加权空间似然编码,联合建模语音位置与分离掩码
- 在相同设置下,定位与分离性能均显著优于基线方法
- 可替代独立定位系统,适合高要求语音处理场景
由于鲁棒性和灵活性,神经网络驱动的波束成形器在存在噪声、混响及多说话人干扰的复杂环境下广受欢迎。可通过时频掩码和相对于固定空间网格的说话人相对方向来估计波束成形器参数。在一定程度上,通过增加空间分区数量超过语音源数量,实现说话人无关性。本文分析如何将掩码与定位信息共同编码至该网格,以实现两者的联合估计。提出掩码加权空间似然编码方法,实验表明其在定位与掩码估计两项任务中均显著优于分别针对单一任务优化的基线编码方式。在相同设置下,联合估计表现更优。最终,提出一种通用方案,仅需调整训练框架即可取代上游声源定位系统,在性能敏感场景中具有重要应用价值。
原文摘要 · Abstract (English)
Due to their robustness and flexibility, neural-driven beamformers are a popular choice for speech separation in challenging environments with a varying amount of simultaneous speakers alongside noise and reverberation. Time-frequency masks and relative directions of the speakers regarding a fixed spatial grid can be used to estimate the beamformer's parameters. To some degree, speaker-independence is achieved by ensuring a greater amount of spatial partitions than speech sources. In this work, we analyze how to encode both mask and positioning into such a grid to enable joint estimation of both quantities. We propose mask-weighted spatial likelihood coding and show that it achieves considerable performance in both tasks compared to baseline encodings optimized for either localization or mask estimation. In the same setup, we demonstrate superiority for joint estimation of both quantities. Conclusively, we propose a universal approach which can replace an upstream sound source localization system solely by adapting the training framework, making it highly relevant in performance-critical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。