构建多模态对话识别数据集,解决多人混杂说话的听觉难题
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
- 融合音视频与上下文线索,识别谁在何时说什么
- 语音重叠达100%,纯音频错误率超100%
- 视觉信息使识别性能提升50%,适合多模态研究者
我们提出了第九届CHiME挑战赛中的多模态上下文感知识别(MCoRec)任务,旨在通过音频、视觉和上下文线索解决单房间内重叠对话的鸡尾酒会问题。该任务捕捉自然的多人对话场景,录音聚焦非剧本化、随意的群体聊天,导致高达100%的语音重叠和高度碎片化的发言轮次。系统需从音视频记录中联合转录每位说话人的话语,并将他们聚类至各自对话中,以回答“谁在何时、说什么、与谁说?”的问题。纯音频基线系统词错误率超过100%,而引入视觉线索后性能大幅提升50%,凸显多模态融合的重要性。本文介绍了任务动机、数据采集流程,并报告了为MCoRec开发的基线系统。
原文摘要 · Abstract (English)
We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a single-room setting using audio, visual, and contextual cues. MCoRec captures natural multi-party conversations where the recordings focus on unscripted, casual group chats, leading to extreme speech overlap of up to 100% and highly fragmented conversational turns. The task requires systems to answer the question "Who speaks when, what, and with whom?" by jointly transcribing each speaker's speech and clustering them into their respective conversations from audio-visual recordings. Audio-only baselines exceed 100% word error rate, whereas incorporating visual cues yields substantial 50% improvements, highlighting the importance of multi-modality. In this manuscript, we present the motivation behind the task, outline the data collection process, and report the baseline systems developed for the MCoRec.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。