arXiv:2601.02231eess.AS2026-01中稿 · HSCMA 2026被引 1

探索空间特征对语音分离模型的提升效果

On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization

  • 在单通道模型中引入多通道空间特征进行条件控制
  • 会议类数据集上性能略有提升,但增幅有限
  • 提示基础模型已充分提取说话人关键信息

近期语音分离技术利用WavLM等大型预训练基础模型,在多个数据集上取得领先性能。现有系统如DiariZen依赖单通道音频表示,无法利用多通道录音中的空间线索。本文通过多种策略将多通道空间特征引入当前最先进的单通道分离系统,评估其对性能的影响。在会议风格数据集上的实验表明,空间信息能提升分离效果,但整体增益小于预期,说明WavLM各层聚合特征已包含大量区分说话人的必要信息,即使在重叠语音区域也具备足够判别力。该结果揭示了空间线索在基础模型驱动语音分离中的潜力与局限。

原文摘要 · Abstract (English)

Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZen leverage these rich single-channel representations, but are limited to single-channel audio, preventing the use of spatial cues available in multi-channel recordings. This work analyzes the impact of incorporating spatial information into a state-of-the-art single-channel diarization system by evaluating several strategies for conditioning the model on multi-channel spatial features. Experiments on meeting-style datasets indicate that spatial information can improve diarization performance, but the overall improvement is smaller than expected for the proposed system, suggesting that the features aggregated over all WavLM layers already capture much of the information needed for accurate speaker discrimination, also in overlapping speech regions. These findings provide insight into the potential and limitations of using spatial cues to enhance foundation model-based diarization.

语音分离空间特征基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。