arXiv:2607.12647eess.AS2026-07中稿 · IWAENC 2026

将空间信息显式融入语音模型,提升语音分离效果

Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization

  • 用显式空间特征条件化下游网络,替代传统波束成形器
  • 在重叠语音区域,波束成形器反而降低性能,最佳结果来自特征条件化
  • 适合做语音分离与会议处理的研究者参考

多通道输入提取的空间信息已被证明可提升会议处理任务(如说话人分离)的性能。同时,基于大预训练单通道基础模型(如 WavLM)提取的特征,已实现最先进的说话人分离效果。本文比较了三种将空间特征融入基础模型驱动的说话人分离系统的方法:波束成形器与单通道基础模型的级联、多通道基础模型,以及对下游网络显式地以提取的空间特征进行条件化。结果表明,在重叠语音区域,波束成形器前端甚至会损害分离性能;而采用条件化方式取得最优效果,证明显式引入空间特征是支持基础模型说话人分离的一种有竞争力的方法。进一步的误差分析显示,该方法显著减少了仅使用谱特征或仅使用空间特征时产生的错误。

原文摘要 · Abstract (English)

Spatial information gleaned from multi-channel input has been shown to lead to improvements in meeting processing tasks like diarization and source separation. At the same time, diarization based on features extracted by large pretrained single-channel foundation models, such as WavLM, achieved state-of-the-art performance. This work compares three approaches to integrate spatial features into foundation model-based diarization systems: the cascade of a beamformer and a single-channel foundation model, a multi-channel foundation model, and the conditioning of the downstream network on explicitly extracted spatial features. Results show that the beamformer front-end is even detrimental to diarization performance in regions of overlapped speech, while best performance is achieved with the conditioning, demonstrating that the incorporation of explicit spatial features is a competitive approach to foundation-model-supported diarization. This approach is further subjected to a detailed error analysis showing that the conditioning system removes errors to a good extent that would occur when either only spectral or only spatial features were used.

说话人分离空间信息基础模型语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。