将LeJEPA扩展至音视频自监督学习,无需解码器和对比负样本。
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

- 采用早期融合ViT与模态丢弃掩码,对齐全局与局部视图嵌入。
- 在VGGSound上达到57.1%准确率,在AudioSet上达32.7 mAP。
- 架构简洁,支持零样本音视频检索,适合多模态预训练研究。
我们提出AV-JEPA,是LeJEPA在音视频自监督学习中的优雅扩展。通过早期融合的Vision Transformer与模态丢弃作为掩码,模型训练目标是使全局视图与各模态局部视图的嵌入对齐,同时使用SIGReg目标鼓励理论上最优的分布。该方法在隐空间实现跨模态对齐,具有极简架构:无解码器、无EMA教师、无复杂多目标损失、无需对比负样本。所提出的AV-JEPA主干在VGGSound数据集上达到57.1% top-1分类准确率,在AudioSet上达到32.7 mAP,且可直接支持零样本音视频检索。
原文摘要 · Abstract (English)
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no decoder, EMA teacher, complex multi-term losses, or contrastive negatives. The proposed AV-JEPA backbone delivers competitive classification performance on VGGSound (57.1% top-1) and AudioSet (32.7 mAP) and supports zero-shot audio-video retrieval out of the box.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。