arXiv:2604.04229cs.MMcs.AI2026-04中稿 · IEEE ICME 2026

通过三层语义关联机制,实现无监督音视频表征学习的精准对齐。

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning

  • 分层设计三重语义关联:全局几何、局部邻域、样本级条件保持。
  • 在AVE和VEGAS数据集上,mAP显著优于现有无监督基线。
  • 适合做音视频多模态表征学习的研究者或工程师参考。

从弱配对、无标签的音视频数据中学习对齐的多模态嵌入极具挑战:现有方法通常仅提供预提取特征,视频片段包含多个事件,且存在虚假共现。我们提出HSC-MAE(层次语义相关性感知掩码自编码器),一种双路径师生框架,强制在三个互补层级——从粗到细——实现语义一致性:(i) 全局级通过DCCA实现模态不变子空间内的标准几何相关性;(ii) 局部级通过教师挖掘的软top-k亲和度保留语义相似实例间的多正例关系结构;(iii) 样本级通过掩码自编码确保在部分观测下嵌入仍保留判别性语义内容。具体地,学生路径采用掩码特征重建与亲和加权软top-k InfoNCE训练;基于未掩码输入的EMA教师通过CCA路径提供稳定的几何结构与软正例。可学习的多任务权重协调冲突目标,可选的蒸馏损失将教师几何信息传递至学生。在AVE和VEGAS上的实验表明,相较于强基线,HSC-MAE显著提升mAP,验证其能生成鲁棒且结构良好的音视频表示。

原文摘要 · Abstract (English)

Learning aligned multimodal embeddings from weakly paired, label-free corpora is challenging: pipelines often provide only pre-extracted features, clips contain multiple events, and spurious co-occurrences. We propose HSC-MAE (Hierarchical Semantic Correlation-Aware Masked Autoencoder), a dual-path teacher-student framework that enforces semantic consistency across three complementary levels of representation - from coarse to fine: (i) global-level canonical-geometry correlation via DCCA, which aligns audio and visual embeddings within a shared modality-invariant subspace; (ii) local-level neighborhood-semantics correlation via teacher-mined soft top-k affinities, which preserves multi-positive relational structure among semantically similar instances; and (iii) sample-level conditional-sufficiency correlation via masked autoencoding, which ensures individual embeddings retain discriminative semantic content under partial observation. Concretely, a student MAE path is trained with masked feature reconstruction and affinity-weighted soft top-k InfoNCE; an EMA teacher operating on unmasked inputs via the CCA path supplies stable canonical geometry and soft positives. Learnable multi-task weights reconcile competing objectives, and an optional distillation loss transfers teacher geometry into the student. Experiments on AVE and VEGAS demonstrate substantial mAP improvements over strong unsupervised baselines, validating that HSC-MAE yields robust and well-structured audio-visual representations.

多模态学习自编码器无监督音视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。