用视觉文本提示增强视频情感识别的多模态理解能力
Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
- 将空间标注与上下文线索融合为统一提示策略
- 在多个数据集上实现显著优于现有方法的零样本识别效果
- 适合需要精准捕捉非语言情绪线索的研究与应用
视觉大语言模型(VLLMs)在多模态理解方面展现出巨大潜力,但在基于视频的情感识别应用中仍受限于空间与上下文感知不足。传统方法侧重孤立面部特征,常忽略肢体语言、环境背景及社会互动等关键非语言线索,导致真实场景下鲁棒性下降。为此,我们提出集合视觉-文本提示(SoVTP),通过整合空间标注(如边界框、面部关键点)、生理信号(面部动作单元)以及上下文线索(身体姿态、场景动态、他人情绪),构建统一的提示机制。SoVTP在保留整体场景信息的同时,支持对面部肌肉运动与人际互动的细粒度分析。大量实验表明,相比现有视觉提示方法,SoVTP在多个数据集上均实现显著提升,验证了其在增强VLLMs视频情感识别能力方面的有效性。
原文摘要 · Abstract (English)
Vision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. Traditional approaches, which prioritize isolated facial features, often neglect critical non-verbal cues such as body language, environmental context, and social interactions, leading to reduced robustness in real-world scenarios. To address this gap, we propose Set-of-Vision-Text Prompting (SoVTP), a novel framework that enhances zero-shot emotion recognition by integrating spatial annotations (e.g., bounding boxes, facial landmarks), physiological signals (facial action units), and contextual cues (body posture, scene dynamics, others' emotions) into a unified prompting strategy. SoVTP preserves holistic scene information while enabling fine-grained analysis of facial muscle movements and interpersonal dynamics. Extensive experiments show that SoVTP achieves substantial improvements over existing visual prompting methods, demonstrating its effectiveness in enhancing VLLMs' video emotion recognition capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。