一招解决会议语音分离与说话人识别,支持实时分段处理。
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
- 用联合统计模型同时完成语音分离和说话人识别
- 在LibriCSS数据集上词错误率优于传统分步方法
- 可自动估算每段活跃人数,实现在线块处理
本文提出一种会议语音的同步说话人分离与识别方法。该方法结合复数角中心高斯混合模型(cACGMM)用于语音源分离,以及冯·米塞斯-费舍尔混合模型(VMFMM)用于说话人识别,构建统一的统计框架,同时利用空间与频谱信息。为支持分段处理,提出一种在段落级别估计活跃说话人数量的方法,解决了跨段语音标签混淆问题,使系统具备块级在线处理潜力。在LibriCSS会议语料库上的实验表明,该集成方法在每段及每会话层面的词错误率(WER)均优于传统的先识别后增强的级联方法。
原文摘要 · Abstract (English)
We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (VMFMM) for diarization in a joint statistical framework. Through the integration, both spatial and spectral information are exploited for diarization and separation. We also develop a method for counting the number of active speakers in a segment of a meeting to support block-wise processing. While the total number of speakers in a meeting may be known, it is usually not known on a per-segment level. With the proposed speaker counting, joint diarization and source separation can be done segment-by-segment, and the permutation problem across segments is solved, thus allowing for block-online processing in the future. Experimental results on the LibriCSS meeting corpus show that the integrated approach outperforms a cascaded approach of diarization and speech enhancement in terms of WER, both on a per-segment and on a per-meeting level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。