arXiv:2602.08309cs.CV2026-02

通过跨模态交互增强,解决音视频学习中的错位问题。

CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment

  • 用双模块动态优化音视频时空对齐,缓解错位干扰。
  • 在多个基准上达领先效果,无需微调主干网络。
  • 适合需要高鲁棒性音视频理解的场景应用。

音视频学习常因离屏声源和背景杂乱导致模态错位,现有方法易放大无关区域或时刻,造成训练不稳定与表征质量下降。为此,我们提出一种新型的Caption-aligned and Agreement-guided Enhancement(CAE-AV)框架,包含两个互补模块:跨模态一致性引导的时空增强(CASTE)与标题对齐的显著性引导增强(CASE),以缓解音视频错位问题。CASTE通过评估帧级音视频一致性,动态平衡空间与时间关系,在错位情况下仍能捕获前后帧的关键信息。CASE在选定时空位置注入跨模态语义引导,利用高层语义线索进一步缓解错位。此外,设计轻量级目标函数:标题-模态InfoNCE、视觉-音频一致性与熵正则化,指导标记选择并强化跨模态语义对齐。在冻结主干网络条件下,CAE-AV在AVE、AVVP、AVS和AVQA等多个基准上达到最优性能,定性分析也验证其对音视频错位的强鲁棒性。

原文摘要 · Abstract (English)

Audio-visual learning suffers from modality misalignment caused by off-screen sources and background clutter, and current methods usually amplify irrelevant regions or moments, leading to unstable training and degraded representation quality. To address this challenge, we proposed a novel Caption-aligned and Agreement-guided Enhancement framework (CAE-AV) for audio-visual learning, which used two complementary modules: Cross-modal Agreement-guided Spatio-Temporal Enrichment (CASTE) and Caption-Aligned Saliency-guided Enrichment (CASE) to relieve audio-visual misalignment. CASTE dynamically balances spatial and temporal relations by evaluating frame-level audio-visual agreement, ensuring that key information is captured from both preceding and subsequent frames under misalignment. CASE injects cross-modal semantic guidance into selected spatio-temporal positions, leveraging high-level semantic cues to further alleviate misalignment. In addition, we design lightweight objectives, caption-to-modality InfoNCE, visual-audio consistency, and entropy regularization to guide token selection and strengthen cross-modal semantic alignment. With frozen backbones, CAE-AV achieves state-of-the-art performance on AVE, AVVP, AVS, and AVQA benchmarks, and qualitative analyses further validate its robustness against audio-visual misalignment.

音视频对齐跨模态学习增强模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。