arXiv:2608.09451eess.AScs.SD2026-08中稿 · MLSP 2026

无需训练的动态聚类方法,提升长语音分离中跨段落排列对齐效果。

Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation

论文配图:Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation
图 1 · 摘自论文原文
  • 基于说话人嵌入参考池,用余弦相似度动态预测跨段排列。
  • 在密集和稀疏场景下均优于现有方法,尤其在长间隙稀疏场景表现更佳。
  • 可直接作为后处理模块使用,对说话人数量估计误差有强鲁棒性。

长语音分离通常采用分段分离拼接范式,将录音划分为短段独立处理再拼接。其核心挑战在于预测跨段落的排列。本文提出一种无需训练的动态聚类方法,利用说话人嵌入参考池实现跨段落排列对齐。该方法通过计算当前段嵌入与参考池中嵌入的余弦相似度来预测排列,并根据整体余弦相似度保留最具代表性的说话人嵌入以更新参考池。作为可即插即用的后处理模块,该方法在密集和稀疏长语音场景下均表现优越,尤其在存在长语音间隔的困难稀疏场景中效果显著,且在未知说话人数量情况下对说话人数量估计误差具有强鲁棒性。

原文摘要 · Abstract (English)

Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its challenge lies in predicting cross-segment permutations. This paper proposes a training-free dynamic clustering approach for cross-segment permutation alignment using speaker embedding reference pools. The method predicts the permutation using the cosine similarity between current segment embeddings and the reference pools. The approach updates reference pools by retaining the most representative speaker embeddings based on their overall cosine similarity with existing references. As a plug-and-play post-processing module compatible with existing separation models, the proposed method demonstrates superior performance compared to existing methods on dense and sparse long speech scenarios, particularly in challenging sparse scenarios with extended utterance gaps, and further shows robustness to speaker count estimation errors in unknown speaker count scenarios.

语音分离动态聚类说话人嵌入长语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。