arXiv:2507.18161eess.AScs.CL2025-07综述被引 8

综述CHiME-7/8远场语音识别挑战,揭示端到端主流与分离技术瓶颈。

Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges

  • 采用端到端ASR系统,依赖大规模预训练模型降低数据需求。
  • 仍依赖引导式语音分离,神经分离技术在复杂场景下表现不足。
  • 通过目标说话人重分角色提升准确率,精确数人是关键。

CHiME-7和8远场语音识别(DASR)挑战聚焦于多通道、可泛化、联合自动语音识别(ASR)与说话人辨识的对话语音处理。共有9支团队提交32种不同系统,推动了该领域的前沿研究。本文概述挑战的设计、评估指标、数据集与基线系统,并分析参赛方案的关键趋势:1)多数参与者采用端到端(e2e)ASR系统,而混合系统在以往挑战中更常见,此转变主要得益于鲁棒的大规模预训练模型,降低了e2e-ASR的数据负担;2)尽管神经语音分离与增强(SSE)取得进展,所有团队仍严重依赖引导式源分离,表明当前神经SSE技术尚无法可靠应对复杂场景与不同录音设置;3)所有最优系统均通过目标说话人辨识技术进行辨识精炼,首次辨识中的准确说话人数至关重要,以避免误差累积,CHiME-8参与者尤其关注此环节;4)下游会议摘要任务与转录质量的相关性较弱,因大语言模型对错误具有强容错能力,在NOTSOFAR-1场景下,即使时间受限最小排列词错误率超过50%的系统,其表现仍与最佳系统(约11%)大致相当;5)尽管有进步,使用计算量大的系统集成,准确转录挑战性声学环境下的自发对话仍困难重重。

原文摘要 · Abstract (English)

The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse systems, these challenges have contributed to state-of-the-art research in the field. This paper outlines the challenges' design, evaluation metrics, datasets, and baseline systems while analyzing key trends from participant submissions. From this analysis it emerges that: 1) Most participants use end-to-end (e2e) ASR systems, whereas hybrid systems were prevalent in previous CHiME challenges. This transition is mainly due to the availability of robust large-scale pre-trained models, which lowers the data burden for e2e-ASR. 2) Despite recent advances in neural speech separation and enhancement (SSE), all teams still heavily rely on guided source separation, suggesting that current neural SSE techniques are still unable to reliably deal with complex scenarios and different recording setups. 3) All best systems employ diarization refinement via target-speaker diarization techniques. Accurate speaker counting in the first diarization pass is thus crucial to avoid compounding errors and CHiME-8 DASR participants especially focused on this part. 4) Downstream evaluation via meeting summarization can correlate weakly with transcription quality due to the remarkable effectiveness of large-language models in handling errors. On the NOTSOFAR-1 scenario, even systems with over 50% time-constrained minimum permutation WER can perform roughly on par with the most effective ones (around 11%). 5) Despite recent progress, accurately transcribing spontaneous speech in challenging acoustic environments remains difficult, even when using computationally intensive system ensembles.

语音识别远场语音说话人辨识CHiME挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。