arXiv:2506.01916eess.AS2025-06ACL被引 6

端到端联合训练,实现多说话人会议的说话人归属语音识别

DNCASR: End-to-End Training for Speaker-Attributed ASR

  • 双编码器分别提取说话人特征和语音波形信息,通过关联解码器统一建模
  • 在AMI-MDM数据集上,说话人归属词错误率降低9.0%(相对)
  • 适合处理真实会议中重叠语音场景,对多说话人转录有实用价值

本文提出DNCASR,一种可端到端训练的新型系统,用于联合神经说话人聚类与自动语音识别(ASR),实现长时多说话人会议的说话人归属转录。DNCASR采用两个独立编码器分别编码全局说话人特征与局部波形信息,并配备两个关联解码器生成说话人归属转录结果。通过关联解码器,整个系统可在统一损失函数下联合训练。采用序列化训练策略,有效缓解真实会议中重叠语音问题,关联机制提升了重叠段落中的说话人索引预测准确率。在AMI-MDM会议语料库上的实验表明,联合训练的DNCASR优于无解码器关联的并行系统。使用cpWER衡量说话人归属词错误率,在AMI-MDM评测集上实现9.0%的相对下降。

原文摘要 · Abstract (English)

This paper introduces DNCASR, a novel end-to-end trainable system designed for joint neural speaker clustering and automatic speech recognition (ASR), enabling speaker-attributed transcription of long multi-party meetings. DNCASR uses two separate encoders to independently encode global speaker characteristics and local waveform information, along with two linked decoders to generate speaker-attributed transcriptions. The use of linked decoders allows the entire system to be jointly trained under a unified loss function. By employing a serialised training approach, DNCASR effectively addresses overlapping speech in real-world meetings, where the link improves the prediction of speaker indices in overlapping segments. Experiments on the AMI-MDM meeting corpus demonstrate that the jointly trained DNCASR outperforms a parallel system that does not have links between the speaker and ASR decoders. Using cpWER to measure the speaker-attributed word error rate, DNCASR achieves a 9.0% relative reduction on the AMI-MDM Eval set.

语音识别说话人归属端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。