arXiv:2502.12454cs.CVcs.AI2025-02中稿 · MRAC'25: 3rd Inter…被引 7

用大模型零样本标注人脸情绪,成本低且可行

Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches

  • 用GPT-4o-mini零样本识别面部情绪,不需训练
  • 七分类准确率约50%,三分类达64%
  • 多帧融合小幅提升效果,适合低成本场景

本研究探讨了使用大型多模态模型(LMMs)在日常场景中自动标注人类情绪的可行性与表现。我们在公开数据集FERV39k的DailyLife子集上进行实验,采用GPT-4o-mini模型对视频片段提取的关键帧进行快速零样本标注。在七类情绪分类(“愤怒”、“厌恶”、“恐惧”、“快乐”、“中性”、“悲伤”、“惊讶”)下,平均精确度约为50%;当限制为三类情绪分类(负面/中性/正面)时,平均精确度提升至约64%。此外,我们探索了在1-2秒视频片段内整合多帧以提升标注性能并降低耗时的方法,结果表明该策略可略微提高标注准确率。总体而言,初步结果表明零样本LMM在人脸情绪标注任务中具有应用潜力,为降低标注成本、拓展LMM在复杂多模态环境中的适用性提供了新路径。

原文摘要 · Abstract (English)

This study investigates the feasibility and performance of using large multimodal models (LMMs) to automatically annotate human emotions in everyday scenarios. We conducted experiments on the DailyLife subset of the publicly available FERV39k dataset, employing the GPT-4o-mini model for rapid, zero-shot labeling of key frames extracted from video segments. Under a seven-class emotion taxonomy ("Angry," "Disgust," "Fear," "Happy," "Neutral," "Sad," "Surprise"), the LMM achieved an average precision of approximately 50%. In contrast, when limited to ternary emotion classification (negative/neutral/positive), the average precision increased to approximately 64%. Additionally, we explored a strategy that integrates multiple frames within 1-2 second video clips to enhance labeling performance and reduce costs. The results indicate that this approach can slightly improve annotation accuracy. Overall, our preliminary findings highlight the potential application of zero-shot LMMs in human facial emotion annotation tasks, offering new avenues for reducing labeling costs and broadening the applicability of LMMs in complex multimodal environments.

情绪识别零样本多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。