arXiv:2505.14433eess.AScs.SD2025-05中稿 · Eusipco 2025被引 1

利用距离与房间信息提升单通道语音分离效果

Single-Channel Target Speech Extraction Utilizing Distance and Room Clues

  • 在时频域构建可学习的距离与房间嵌入模型
  • 在模拟与真实数据集上实现更优的语音分离性能
  • 适合需要跨房间泛化的语音增强场景

本文旨在利用距离线索和房间信息实现封闭空间内的单通道目标语音提取(TSE)。已有研究证实,距离线索可通过声源的直达/混响比(DRR)反映声源位置信息,可用于语音分离与TSE系统。然而,该线索受房间声学特性(如尺寸、混响时间)显著影响,导致仅依赖距离线索的TSE系统难以在不同房间间泛化。为此,本文提出在基于距离的TSE中引入房间环境信息(房间尺寸与混响时间),以提升泛化能力。特别地,设计了一种在时频域的、结合距离与环境信息的可学习嵌入模型。在模拟与真实采集数据集上的实验结果验证了该方法的有效性。演示材料见 https://runwushi.github.io/distance-room-demo-page/

原文摘要 · Abstract (English)

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/.

语音分离距离线索房间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。