仅用距离信息实现单麦克风语音分离,无需说话人生理特征。
Distance Based Single-Channel Target Speech Extraction
- 利用距离信息与时频单元融合,设计新型分离模型。
- 单房间和多房间场景下均验证有效,可准确估计说话人距离。
- 适合语音增强、智能会议等需要定位的实时应用。
本文旨在仅利用距离信息,在封闭空间中实现单通道目标语音提取(TSE)。这是首个不依赖说话人生理信息、仅使用距离线索完成单通道TSE的研究。受近期基于距离的单通道分离方法启发,我们提出一种新模型,高效融合距离信息与时频(TF)单元,用于目标语音提取。在单房间和多房间场景下的实验结果证明了该方法的可行性和有效性。此外,该方法还可用于混合语音中不同说话人的距离估计。在线演示可访问 https://runwushi.github.io/distance-demo-page。
原文摘要 · Abstract (English)
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for single-channel TSE. Inspired by recent single-channel Distance-based separation and extraction methods, we introduce a novel model that efficiently fuses distance information with time-frequency (TF) bins for TSE. Experimental results in both single-room and multi-room scenarios demonstrate the feasibility and effectiveness of our approach. This method can also be employed to estimate the distances of different speakers in mixed speech. Online demos are available at https://runwushi.github.io/distance-demo-page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。