arXiv:2505.13971cs.SDcs.AI2025-05中稿 · Interspeech 2025被引 10

多模态会议语音处理挑战赛,提升音视频说话人分离与识别效果

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

  • 融合音视频信息进行多设备会议转录
  • 最佳模型在说话人分离上误差率降7.43%
  • 适合语音识别与人机交互研究者参考

会议是语音应用的重要场景,但复杂的声学条件带来挑战。本文总结了2025年在Interspeech会议上举办的MISP挑战赛成果,聚焦多模态、多设备会议转录,结合音频与视频模态。任务包括音视频说话人分离(AVSD)、音视频语音识别(AVSR)以及音视频分离与识别(AVDR)。文中介绍了挑战的目标、任务设计、数据集、基线系统及参赛者提出的方案。最佳系统显著优于基线:最优AVSD模型达到8.09%的分离错误率(DER),改善7.43%;最优AVSR系统取得9.48%的字符错误率(CER),改进10.62%;最佳AVDR系统实现11.56%的拼接最小排列字符错误率(cpCER),提升72.49%。

原文摘要 · Abstract (English)

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.

音视频融合说话人分离会议转录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。