arXiv:2505.16387eess.AS2025-05中稿 · Interspeech2025被引 3

用多通道音频优化语音分角色,夺冠系统精度达8.09%。

Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge

  • 先用单通道数据生成初始分角色结果
  • 再融合多通道信息提升精度,最终DER为8.09%
  • 适合需要高精度语音分离的会议与访谈场景

本文介绍了为2025年多模态语音处理挑战赛(MISP 2025 Challenge)开发的说话人分角色系统。首先,采用序列到序列神经分角色(S2SND)框架,基于单通道音频生成初始预测结果;随后,将原始S2SND框架扩展为多通道序列到序列神经分角色(MC-S2SND),利用多通道音频进一步优化初始结果。最终系统在比赛数据库的评估集上取得8.09%的分角色错误率(DER),在说话人分角色任务中排名第一。

原文摘要 · Abstract (English)

This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequence Neural Diarization (S2SND) framework to generate initial predictions using single-channel audio. Then, we extend the original S2SND framework to create a new version, Multi-Channel Sequence-to-Sequence Neural Diarization (MC-S2SND), which refines the initial results using multi-channel audio. The final system achieves a diarization error rate (DER) of 8.09% on the evaluation set of the competition database, ranking first place in the speaker diarization task of the MISP 2025 Challenge.

语音分角色多通道音频序列模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。