用自验证方法低成本生成高质量表情指令数据,提升视觉大模型识情绪能力。
Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation
- 融合人类先验知识与多粒度情绪关联,生成统一标注。
- 自验证策略结合不确定性采样,提升预测准确率32.7%以上。
- 适合需要高精度表情分析的AI交互、医疗诊断场景。
视觉大语言模型(VLLM)在面部情绪感知中对实现自然人机交互至关重要。然而,粗粒度与细粒度面部情绪分析的高质量标注需耗费大量专业人力。缺乏此类高质量指令数据限制了VLLM在情绪感知中的表现。为此,我们提出一种带情绪知识增强的自验证方法(SEKE),利用闭源VLLM低成本生成多粒度情绪分析的高质量指令数据。该方法通过整合人类先验知识,基于三种情绪描述层级(离散表情、效价-唤醒度、动作单元)间的内在关联,可靠生成全面标注。进一步引入不确定性感知蒙特卡洛采样自验证策略(SV-UAMC),高效提取更准确的VLLM预测,提升标注可靠性。最终构建包含三类完整描述的面部情绪指令数据集(FEID),提供粗粒度与细粒度情感信息,用于有效模型训练。此外,我们引入面部情绪分析基准(FEAB)以评估VLLM相应能力。本方法在三个下游任务上显著优于现有最先进方法。
原文摘要 · Abstract (English)
Facial emotion perception in the vision large language model (VLLM) is crucial for achieving natural human-machine interaction. However, creating high-quality annotations for both coarse- and fine-grained facial emotion analysis demands costly expertise. The lack of such high-quality instruction data limits the performance of VLLMs in facial emotion perception. To address this, we propose a self-verification approach with emotion knowledge enhancement (SEKE), which generates high-quality instruction data for multi-grained emotion analysis cost-effectively using closed-source VLLM. This approach integrates prior human knowledge to VLLM inference, guided by the inherent correlations between three grained levels of emotion descriptions, i.e., discrete expression, valence-arousal, and action unit, to reliably generate comprehensive annotations. A self-verification strategy with Uncertainty-Aware Monte Carlo sampling (SV-UAMC) is further embedded to efficiently extract more accurate VLLM predictions, further improving annotation reliability. Consequently, we construct a facial emotion instruction dataset (FEID) containing three comprehensive descriptions, which provides coarse- and fine-grained emotional information for effective model training. Additionally, we introduce a facial emotion analysis benchmark (FEAB) to measure the VLLM's corresponding ability. Our method significantly outperforms state-of-the-art methods on three downstream facial emotion analysis tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。