arXiv:2603.01415eess.AS2026-03

用视频+音频识别多人同时对话,准确率创新高。

The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge

  • 融合360度视频与单通道音频,分步处理说话人检测、语音提取和识别。
  • 在开发集上达到31.40%的说话人词错误率,聚类F1达1.0。
  • 适合研究多说话人语音识别与社交场景下对话分离的学者。

本文介绍我们针对CHiME-9 MCoRec挑战赛的提交方案,该挑战聚焦于室内社交环境中多个并发自然对话的识别与聚类。与传统以单一主题为中心的会议不同,此场景包含最多八名说话人、四组并行对话,语音重叠率超过90%。为此,我们提出一种多模态级联系统,利用同步的360度视频提取的逐说话人视觉流,结合单通道音频。通过增强的音视频预训练模型,改进了流水线中的三个组件:主动说话人检测(ASD)、音视频目标语音提取(AVTSE)和音视频语音识别(AVSR)。AVSR模块进一步引入Whisper和大语言模型(LLM)技术提升转录准确率。最佳单一级联系统在开发集上的说话人词错误率(WER)为32.44%。通过ROVER融合不同前端与后端变体输出,将说话人WER降至31.40%。值得注意的是,基于LLM的零样本对话聚类实现1.0的说话人聚类F1分数,最终联合ASR-聚类错误率(JACER)为15.70%。

原文摘要 · Abstract (English)

This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike conventional meetings centered on a single shared topic, this scenario contains multiple parallel dialogues--up to eight speakers across up to four simultaneous conversations--with a speech overlap rate exceeding 90%. To tackle this, we propose a multimodal cascaded system that leverages per-speaker visual streams extracted from synchronized 360 degree video together with single-channel audio. Our system improves three components of the pipeline by leveraging enhanced audio-visual pretrained models: Active Speaker Detection (ASD), Audio-Visual Target Speech Extraction (AVTSE), and Audio-Visual Speech Recognition (AVSR). The AVSR module further incorporates Whisper and LLM techniques to boost transcription accuracy. Our best single cascaded system achieves a Speaker Word Error Rate (WER) of 32.44% on the development set. By further applying ROVER to fuse outputs from diverse front-end and back-end variants, we reduce Speaker WER to 31.40%. Notably, our LLM-based zero-shot conversational clustering achieves a speaker clustering F1 score of 1.0, yielding a final Joint ASR-Clustering Error Rate (JACER) of 15.70%.

多说话人识别音视频融合对话聚类语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。