arXiv:2607.03018cs.CVcs.SD2026-07

通过多层级一致性约束,提升语音视频说话人检测在噪声等干扰下的鲁棒性。

$C^3$ASD: Multi-Level Consistency-Driven Representation Learning

论文配图:$C^3$ASD: Multi-Level Consistency-Driven Representation Learning
图 1 · 摘自论文原文
  • 引入嵌入层、序列层、预测层三重一致性约束,强化跨模态语义对齐
  • 在多种音频、视觉及联合噪声下准确率提升显著,干净数据上保持竞争力
  • 适合需要高鲁棒性的实际场景应用,如嘈杂环境中的视频监控

主动说话人检测旨在判断视频中可见人物是否正在讲话。尽管现有音视频融合方法在干净数据上表现良好,但在真实世界干扰(如背景噪声、遮挡或双模态同时退化)下性能下降。我们归因于缺乏显式的一致性约束,导致模型难以学习到鲁棒且语义对齐的跨模态表示。模型易依赖脆弱的模态特异性捷径,在干扰条件下失效。为此,我们提出 $C^3$ASD,一种多层次一致性驱动框架,包含三项互补约束:嵌入层跨模态一致性,在语音时对齐音视频表征;序列层单模态一致性,通过轨迹感知对比学习分离说话与非说话聚类;预测层一致性,借助知识蒸馏稳定融合结果。大量实验表明,该方法在多种音频、视觉及联合噪声下均有显著提升,同时在干净数据上保持竞争力。

原文摘要 · Abstract (English)

Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as background noise, occlusion, or simultaneous modality degradation. We attribute this limitation to the absence of explicit consistency constraints that promote robust, semantically aligned representations across modalities. Without such guidance, models tend to learn fragile modality-specific shortcuts that fail under corrupted conditions. We propose $C^3$ASD, a multi-level consistency-driven framework with three complementary constraints: embedding-level inter-modality consistency aligns audio-visual representations during speech; sequence-level intra-modality consistency separates speaking and non-speaking clusters via track-aware contrastive learning; and prediction-level consistency stabilizes fusion through knowledge distillation. Extensive experiments demonstrate significant improvements under diverse audio, visual and joint corruptions, while maintaining competitive performance on clean data.

说话人检测音视频融合鲁棒性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。