arXiv:2410.02244cs.CV2024-10EMNLP被引 19

用视觉提示提升大模型识情绪能力,精准定位人脸表情。

Visual Prompting in LLMs for Enhancing Emotion Recognition

  • 引入视觉提示集(SoV),利用框和关键点精确定位人脸
  • 零样本下提升人脸计数与情绪分类准确率
  • 适合需要高精度情绪分析的智能交互场景

视觉大语言模型(VLLMs)正在重塑计算机视觉与自然语言处理的交叉领域。然而,利用视觉提示增强情绪识别的潜力仍鲜被探索。传统VLLM方法在空间定位上表现不佳,常丢失全局上下文信息。为此,我们提出一种视觉提示集(SoV)方法,通过边界框和面部关键点等空间信息,精准标记目标,从而提升零样本情绪识别性能。SoV在保持丰富图像上下文的同时,提高了人脸数量估计和情绪类别划分的准确性。通过对最新商业或开源VLLM的多轮实验与分析,验证了该方法在真实环境下的有效性。结果表明,将空间视觉提示融入VLLM可显著提升情绪识别性能。

原文摘要 · Abstract (English)

Vision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing. Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped. Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context. To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely. SoV improves accuracy in face count and emotion categorization while preserving the enriched image context. Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments. Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance.

情绪识别视觉提示大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。