构建跨模态统一表征,实现音视频图文联合理解与检索。
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- 采用大规模对比学习,融合音视频图文三模态数据训练统一编码器。
- 在1亿级音视频对上生成高质量字幕,支持多任务零样本性能提升。
- 适配声音事件检测等细粒度任务,适合多模态感知研究者使用。
我们提出感知编码器音视频(PE-AV),一种基于缩放对比学习的音频与视频理解编码器家族。基于已有模型PE,PE-AV将表征扩展至音频领域,并原生支持音视频、音频文本、视频文本之间的联合嵌入。其统一的跨模态嵌入使语音检索等新任务成为可能,在标准音视频基准测试中达到新的性能上限。我们通过构建强大的音视频数据引擎,为约1亿个音视频对合成高质量字幕,实现跨模态的一致性监督。音频数据涵盖语音、音乐和通用音效,避免了以往单一领域限制。我们设计了十种成对对比目标,证明扩大跨模态与字幕类型对可增强对齐效果并提升零样本性能。此外,我们进一步开发了PE-A-Frame,通过对PE-AV进行帧级对比学习微调,实现音帧与文本的细粒度对齐,适用于声音事件检测等任务。
原文摘要 · Abstract (English)
We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend representations to audio, and natively support joint embeddings across audio-video, audio-text, and video-text modalities. PE-AV's unified cross-modal embeddings enable novel tasks such as speech retrieval, and set a new state of the art across standard audio and video benchmarks. We unlock this by building a strong audiovisual data engine that synthesizes high-quality captions for O(100M) audio-video pairs, enabling large-scale supervision consistent across modalities. Our audio data includes speech, music, and general sound effects-avoiding single-domain limitations common in prior work. We exploit ten pairwise contrastive objectives, showing that scaling cross-modality and caption-type pairs strengthens alignment and improves zero-shot performance. We further develop PE-A-Frame by fine-tuning PE-AV with frame-level contrastive objectives, enabling fine-grained audio-frame-to-text alignment for tasks such as sound event detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。