用大模型理解复杂表情,提升人机情感交互能力。
Compound Expression Recognition via Large Vision-Language Models
- 分两阶段微调大视觉语言模型,先学基础表情再优化复合表达
- 在RAF-DB上达到先进准确率,在C-EXPR-DB上展现强零样本泛化能力
- 适合做情绪分析、人机交互等真实场景应用的开发者参考
复合表情识别(CER)对于理解人类情感和提升人机交互至关重要。然而,由于面部表情的复杂性以及捕捉细微情感线索的困难,该任务面临挑战。为此,我们提出一种利用大视觉语言模型(LVLMs)的新方法。该方法采用两阶段微调:首先在基础表情数据集上微调预训练的LVLM以建立基础模式;其次在复合表情数据集上进一步优化,以精炼视觉-语言特征交互。该方法在RAF-DB数据集上取得先进准确率,并在C-EXPR-DB数据集上展现出优异的零样本泛化能力,展示了其在情感分析与人机交互中的实际应用潜力。
原文摘要 · Abstract (English)
Compound Expression Recognition (CER) is crucial for understanding human emotions and improving human-computer interaction. However, CER faces challenges due to the complexity of facial expressions and the difficulty of capturing subtle emotional cues. To address these issues, we propose a novel approach leveraging Large Vision-Language Models (LVLMs). Our method employs a two-stage fine-tuning process: first, pre-trained LVLMs are fine-tuned on basic facial expressions to establish foundational patterns; second, the model is further optimized on a compound-expression dataset to refine visual-language feature interactions. Our approach achieves advanced accuracy on the RAF-DB dataset and demonstrates strong zero-shot generalization on the C-EXPR-DB dataset, showcasing its potential for real-world applications in emotion analysis and human-computer interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。