arXiv:2608.08510cs.CL2026-08中稿 · ICMI 2026

分析语音识别在嘈杂聚会场景下的多模态系统表现,发现新突破方向。

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

论文配图:From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
图 1 · 摘自论文原文
  • 通过音视频融合分离目标说话人信号
  • 最佳系统实现57%相对错误率降低
  • 大模型用于对话分组与连贯性增强

人类能自然地在嘈杂环境中专注听某一人讲话,而这一‘鸡尾酒会’场景对语音识别系统仍是重大挑战。CHiME-9 MCoRec任务要求系统从音视频输入中识别多个说话人并转录各自对话。本文分析了多种不同设计思路的系统,其中最优方案达到57%的相对错误率降低。研究识别出三大策略:(1)显式或隐式地进行音视频目标语音分离;(2)提升每个目标说话人的音视频语音识别性能;(3)利用大语言模型将说话人分组并增强对话连贯性。分析表明,这些方向针对的是互补性的失败模式,且高语音重叠并非性能差异的主因,挑战了以往认为重叠是主要难点的普遍假设。

原文摘要 · Abstract (English)

Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This "cocktail party" scenario still presents severe challenges to speech recognition systems. The CHiME-9 MCoRec task provides a testbed where systems must recognize groups of speakers and transcribe each of their conversations from audio-visual input. In this work, we analyze a diverse set of systems, representing different design directions for addressing the cocktail-party scenario, where the best system achieves up to 57% relative error reduction. We identify three main strategies: (1) explicit or implicit audio-visual target speech separation, (2) improved audio-visual speech recognition for each target speaker, and (3) the use of large language models to group speakers into conversations and enhance conversational consistency. Our analysis shows that these directions address complementary failure modes of the cocktail-party problem, and that high speech overlap alone does not explain performance differences, challenging the common assumption that overlap is the primary source of difficulty in cocktail-party recognition.

语音识别多模态鸡尾酒会大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。