arXiv:2608.03964eess.AS2026-08

提出真实双说话人语音提取数据集与模型,精准区分视频中指定说话人。

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

论文配图:Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE
图 1 · 摘自论文原文
  • 用多视角视频+语音联合建模,让模型听懂谁在说话
  • 在真实场景下实现82.2%说话人识别准确率,错误率0.2261
  • 适合做语音对话、会议记录的开发者参考

音频-视觉目标说话人分离需返回视频指示的说话人,但分离器可能忽略视觉线索,重复输出声学主导的声音。本文引入REAL-2MIX,一个包含77个场景、7,598段剪辑(11.84小时)的中文双说话人混合语音真实录制基准,每场景含前置的仅音频A和仅音频B阶段,提供场景内说话人参考。处理后剩余6,038个可评估混合样本和12,076条目标说话人行。此外,从VoxBlink2构建了VOXBLINK2-AVSE,包含250,828对同步音视频-唇部区域图像,覆盖28,421个身份,总计766.17小时语音。所提提取器使用冻结的1,280维投影AV-HuBERT特征,目标条件训练及逐层特征调制。联合评估内容使用Qwen3-ASR-1.7B CER,目标身份使用WeSpeaker ResNet34加重叠语音检测(OSD)。在完整数据集上,最佳模型取得0.2261 CER,82.22%严格输出正确率,69.53%双输出严格成功。

原文摘要 · Abstract (English)

Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant voice. We introduce REAL-2MIX, a Mandarin AV-TSE benchmark of jointly recorded real two-speaker mixtures with synchronized multi-view video. Each scene also contains preceding A-only and B-only stages that provide in-scene speaker references. It contains 77 scenes and 7,598 clips (11.84 hours), including 6,042 dual-annotated mixtures. After processing there leave 6,038 evaluable mixtures and 12,076 target-speaker rows. We additionally curate VOXBLINK2-AVSE from VoxBlink2, comprising 250,828 synchronized audio--lip-ROI pairs from 28,421 identities and 766.17 hours of speech. Our extractor uses frozen, 1,280-dimensional projected AV-HuBERT features, target-conditioned training, and layer-wise feature modulation. We jointly evaluate content with Qwen3-ASR-1.7B CER and target identity with WeSpeaker ResNet34 plus Overlapped Speech Detection (OSD). On the complete manifest, the best archived checkpoint obtains 0.2261 CER, 82.22% strict output correctness, and 69.53% both-output strict success.

语音分离多模态中文数据集说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。