用自监督学习训练多模态人脸语音模型,提升社交情感识别性能
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
- 基于CAV-MAE改进,输入更多帧并自监督预训练于社交数据
- 在情绪识别、笑声检测任务上达到顶尖效果,人格估计也表现优异
- 适合研究多模态感知、情感计算或自监督学习的学者使用
人类社交行为本质上是多模态的,亟需强大的音视频模型进行感知。本文提出Social-MAE,一种基于扩展版对比音视频掩码自编码器(CAV-MAE)的预训练多模态自编码器,其在大规模人类社交交互数据集VoxCeleb2上以自监督方式预训练。我们对CAV-MAE进行了改进,使其可接受更大数量的输入帧。通过微调并在多个社交与情感下游任务中评估,包括情绪识别、笑声检测和显性人格估计,该模型在多模态情绪识别与笑声识别任务上取得当前最优结果,在人格估计任务上也表现良好,验证了领域内自监督预训练的有效性。代码与模型权重已开源。
原文摘要 · Abstract (English)
Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended version of Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE), which is pre-trained on audiovisual social data. Specifically, we modify CAV-MAE to receive a larger number of frames as input and pre-train it on a large dataset of human social interaction (VoxCeleb2) in a self-supervised manner. We demonstrate the effectiveness of this model by finetuning and evaluating the model on different social and affective downstream tasks, namely, emotion recognition, laughter detection and apparent personality estimation. The model achieves state-of-the-art results on multimodal emotion recognition and laughter recognition and competitive results for apparent personality estimation, demonstrating the effectiveness of in-domain self-supervised pre-training. Code and model weight are available here https://github.com/HuBohy/SocialMAE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。