提出可扩展的音视频自编码模型,提升情感面部分析效率与准确性。
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
- 采用双模态掩码策略与增强编码器设计,支持大规模预训练。
- 在17个数据集上实现跨任务最优性能,显著优于现有方法。
- 适合需要高效多模态情感分析的研究者与工业应用开发者。
情感视频面部分析(AVFA)是构建情绪感知智能系统的关键领域,但受限于数据稀缺。近年来,掩码自编码器(MAE)在自监督学习中兴起,并逐步应用于音视频场景。尽管规模扩展在通用多模态学习中被证明至关重要,其对AVFA的具体影响仍不明确。另一个核心挑战是如何通过可扩展的音视频表征捕捉模态内与模态间相关性。为此,我们提出AVF-MAE++,一个面向高效探索AVFA缩放特性的音视频MAE模型家族。该框架引入新颖的音频与视觉双模态掩码策略,并通过更整体的设计增强模态编码器,以更好支持可扩展预训练。此外,提出迭代音视频相关性学习模块,在自监督范式下提升相关性建模能力,弥补先前方法局限。为促进平滑适配并降低过拟合风险,进一步设计渐进语义注入策略,将训练分为三个结构化阶段。在涵盖三大主要任务的17个数据集上的大量实验表明,AVF-MAE++在多个基准上均达到一致领先性能。全面的消融研究验证了各组件的重要性,并深入揭示设计选择背后的机制。代码与模型已公开于GitHub。
原文摘要 · Abstract (English)
Affective video facial analysis (AVFA) has emerged as a key research field for building emotion-aware intelligent systems, yet this field continues to suffer from limited data availability. In recent years, the self-supervised learning (SSL) technique of Masked Autoencoders (MAE) has gained momentum, with growing adaptations in its audio-visual contexts. While scaling has proven essential for breakthroughs in general multi-modal learning domains, its specific impact on AVFA remains largely unexplored. Another core challenge in this field is capturing both intra- and inter-modal correlations through scalable audio-visual representations. To tackle these issues, we propose AVF-MAE++, a family of audio-visual MAE models designed to efficiently investigate the scaling properties in AVFA while enhancing cross-modal correlation modeling. Our framework introduces a novel dual masking strategy across audio and visual modalities and strengthens modality encoders with a more holistic design to better support scalable pre-training. Additionally, we present the Iterative Audio-Visual Correlation Learning Module, which improves correlation learning within the SSL paradigm, bridging the limitations of previous methods. To support smooth adaptation and reduce overfitting risks, we further introduce a progressive semantic injection strategy, organizing the model training into three structured stages. Extensive experiments conducted on 17 datasets, covering three major AVFA tasks, demonstrate that AVF-MAE++ achieves consistent state-of-the-art performance across multiple benchmarks. Comprehensive ablation studies further highlight the importance of each proposed component and provide deeper insights into the design choices driving these improvements. Our code and models have been publicly released at Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。