arXiv:2604.27436eess.ASeess.IV2026-04中稿 · HSCMA 2026 Worksho…被引 1

利用视觉信息提升长对话中多说话人语音识别与分组效果

BUT System Description for CHiME-9 MCoRec Challenge

论文配图:BUT System Description for CHiME-9 MCoRec Challenge
图 1 · 摘自论文原文
  • 用视觉特征引导的长时目标说话人语音识别模型,单次解码处理全程录音
  • 开发集上词错误率33.7%,对话分组F1达0.97,显著优于基线
  • 结合大模型判断语义相似性实现高效对话分组,适合复杂场景语音分析

在大量重叠语音的对话录音中,多说话人自动语音识别仍是未解难题,仅靠音频难以区分目标说话人。视觉线索可缓解歧义,但其在长时音频-视觉(AV)ASR系统中的应用仍受限。CHiME-9 MCoRec任务要求对高度重叠的并行对话音频-视频记录进行转录,并将参与者聚类为对话组。本文提出BUT系统,基于长时目标说话人AV-ASR模型,可在单次解码中处理长时录音。该架构将预训练的NVIDIA Parakeet-v2 ASR模型条件于预训练的AV-HuBERT模型提取的视觉表示。为聚类参与者,采用Qwen3.5-122B大语言模型评估转录内容主题相似性,再通过层次聚类完成分组。在开发集上,系统词错误率为33.7%,聚类F1得分为0.97,相比官方基线分别提升16.2%和0.15绝对值。在评测集上,团队排名第二,词错误率与最佳系统相差0.16%,聚类F1低0.5%。

原文摘要 · Abstract (English)

Multi-talker automatic speech recognition (ASR) in conversational recordings remains an open problem, particularly in scenarios with large portion of overlapping speech where identifying and transcribing a target speaker is difficult from audio alone. Visual cues can help resolve speaker ambiguity, yet their integration into long-context audio-visual (AV) ASR systems has been limited. The CHiME-9 MCoRec task addresses this challenge by requiring transcription of audio-visual recordings of heavily-overlapped parallel conversations, followed by clustering the participants into conversational groups. In this work, we present the BUT system based on a long-context target-speaker AV-ASR model capable of processing long-form recordings in a single decoding pass. Our architecture conditions a pre-trained NVIDIA Parakeet-v2 ASR model on visual representations from a pre-trained AV-HuBERT model. To cluster participants into conversation groups, we employ Qwen3.5-122B LLM to estimate transcript topic similarity followed by hierarchical agglomerative clustering. On the MCoRec development set, the proposed system achieves 33.7% WER and a clustering F1 score of 0.97, improving over the official baseline by 16.2% WER and 0.15 F1 absolute. On the eval set, our team ranked second, being 0.16% WER and 0.5% F1 worse than the best system.

多说话人识别视听语音识别对话聚类大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。