通过属性吸引子与局部依赖建模,提升语音分说话人识别效果
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
- 用多阶段中间表示建模说话人属性,替代直接说话人建模
- 在CALLHOME数据集上性能优于传统方法,提升明显
- 适合需要高精度说话人分离的语音分析场景
近年来,端到端方法在解决多说话人录音中的说话人分段与识别问题上取得了显著进展。其中,编码器-解码器吸引子(EDA)方法能够处理可变数量的说话人,并在训练过程中更有效地引导网络。本文在此基础上扩展吸引子范式,不再直接建模说话人,而是通过多阶段中间表示来捕捉更细致的“说话人属性”。同时,将原架构中的Transformer替换为卷积增强型Transformer(Conformer),以更好地建模局部依赖关系。实验表明,在CALLHOME数据集上,该方法实现了更好的说话人分段性能。
原文摘要 · Abstract (English)
In recent years, end-to-end approaches have made notable progress in addressing the challenge of speaker diarization, which involves segmenting and identifying speakers in multi-talker recordings. One such approach, Encoder-Decoder Attractors (EDA), has been proposed to handle variable speaker counts as well as better guide the network during training. In this study, we extend the attractor paradigm by moving beyond direct speaker modeling and instead focus on representing more detailed `speaker attributes' through a multi-stage process of intermediate representations. Additionally, we enhance the architecture by replacing transformers with conformers, a convolution-augmented transformer, to model local dependencies. Experiments demonstrate improved diarization performance on the CALLHOME dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。