arXiv:2512.24231cs.CVcs.LG2025-12

MotivNet让表情识别模型在真实场景中表现更稳定,无需跨数据集训练。

MotivNet: Evolving Meta-Sapiens into an Emotionally Intelligent Foundation Model

  • 基于Sapiens视觉基础模型,不依赖跨域训练提升泛化能力
  • 在多个数据集上达到顶尖性能,跨域适应性强
  • 适合希望落地真实场景的表情识别研究与应用

本文提出MotivNet,一种具备强泛化能力的通用面部表情识别模型,适用于真实世界应用。当前主流表情识别模型在多样数据上表现不佳,导致实际应用性能下降,阻碍该领域发展。尽管已有研究通过复杂架构改善泛化性,但需跨域训练,与真实场景需求矛盾。MotivNet以Sapiens为骨干网络,通过大规模掩码自编码器预训练获得卓越现实泛化能力,无需跨域训练即可在多数据集上实现竞争力表现。我们定义了基准性能、模型相似性、数据相似性三项标准评估MotivNet作为Sapiens下游任务的可行性。实验表明,MotivNet在多个数据集上优于现有SOTA模型,满足评估标准,验证其有效性,推动表情识别向真实场景应用迈进。代码已开源:https://github.com/OSUPCVLab/EmotionFromFaceImages。

原文摘要 · Abstract (English)

In this paper, we introduce MotivNet, a generalizable facial emotion recognition model for robust real-world application. Current state-of-the-art FER models tend to have weak generalization when tested on diverse data, leading to deteriorated performance in the real world and hindering FER as a research domain. Though researchers have proposed complex architectures to address this generalization issue, they require training cross-domain to obtain generalizable results, which is inherently contradictory for real-world application. Our model, MotivNet, achieves competitive performance across datasets without cross-domain training by using Meta-Sapiens as a backbone. Sapiens is a human vision foundational model with state-of-the-art generalization in the real world through large-scale pretraining of a Masked Autoencoder. We propose MotivNet as an additional downstream task for Sapiens and define three criteria to evaluate MotivNet's viability as a Sapiens task: benchmark performance, model similarity, and data similarity. Throughout this paper, we describe the components of MotivNet, our training approach, and our results showing MotivNet is generalizable across domains. We demonstrate that MotivNet can be benchmarked against existing SOTA models and meets the listed criteria, validating MotivNet as a Sapiens downstream task, and making FER more incentivizing for in-the-wild application. The code is available at https://github.com/OSUPCVLab/EmotionFromFaceImages.

表情识别基础模型泛化能力Sapiens

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。