用统一编码器联合学习音视频,提升表示能力。
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

- 单编码器+跨模态预测,简单高效。
- 音视频互推使双方表征更优,超越单模态基线。
- 零微调下性能领先,适合资源有限场景。
大规模视频数据的自监督学习已成为视觉表征学习的主要范式。由于音频与视觉在视频中自然共现,联合学习两者是合理延伸,但仍有挑战。现有音视频自监督方法依赖模态专用编码器及复杂的对比或重建目标,限制了跨模态协同与可扩展性。联合嵌入预测架构(JEPAs)提供了一种简洁、模态无关的替代方案,但此前主要应用于单一模态。本文提出MJEPA,一种用于音视频学习的联合嵌入预测架构,采用单一统一编码器处理双模态。方法仅使用单一预测目标,同时作用于模态内与跨模态。结果表明跨模态预测至关重要:无此机制时,共享编码器性能低于单模态基线;有之则两模态表征均获益。冻结的ViT-g模型在AudioSet-20K上比最佳先前冻结基线提升超过6.8 mAP,超越全微调模型在ESC-50和FSD50K上的表现,且在视频基准测试中具竞争力,仅使用10倍少的视频数据。
原文摘要 · Abstract (English)
Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in video data, extending this success to jointly learn from both modalities is a natural next step, yet it remains challenging. Existing audio-visual self-supervised methods rely on modality-specific encoders and complex combinations of contrastive or reconstruction objectives, limiting cross-modal synergy and scalability. Joint Embedding Predictive Architectures (JEPAs) offer a simple, modality-agnostic alternative, but have to date been applied primarily to individual modalities. We introduce MJEPA, a joint-embedding predictive architecture for audio-visual learning that uses a single, unified encoder for both modalities. Our approach uses only a single predictive objective, applied both within and across modalities. We show that cross-modal prediction is critical: without it, a shared encoder degrades below unimodal baselines; with it, each modality's representation benefits from the other. Our frozen ViT-g model outperforms the best prior frozen baseline by over 6.8 mAP on AudioSet-20K, surpasses fully finetuned models on ESC-50 and FSD50K, and is competitive on video benchmarks despite using 10x less video data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。