提出可解释多模态情感识别框架,提升对情感线索的感知与推理能力。
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
- 设计视频与音频情感线索桥模块,增强细粒度情感信号捕捉能力。
- 在新构建的EmoCue数据集上,模型情感线索推理性能显著优于基线。
- 提供自动化评估指标与专家标注基准,支持细粒度情感分析研究。
可解释的多模态情感识别在人机交互与社交媒体分析中至关重要。现有方法因两大挑战受限:一是通用模态编码器预训练目标为全局结构与通用语义,难以捕捉细粒度情感线索;二是现有数据集在标注质量与规模间存在权衡,导致情感线索监督不足,限制了线索级推理。此外,现有评估指标难以衡量线索级推理性能。为此,我们提出XEmoGPT,一种具备情感线索感知与推理能力的新框架。其包含视频情感线索桥(VECB)与音频情感线索桥(AECB),通过定制任务增强视频与音频编码器对细粒度情感线索的感知。为进一步支持线索级推理,我们构建大规模数据集EmoCue,用于训练模型进行多模态情感线索推理。同时,提出EmoCue-360自动化评估指标,基于语义相似性提取与匹配情感线索,并发布包含400个专家标注样本的EmoCue-Eval基准。实验表明,XEmoGPT在情感线索感知与推理方面均表现优异。
原文摘要 · Abstract (English)
Explainable Multimodal Emotion Recognition plays a crucial role in applications such as human-computer interaction and social media analytics. However, current approaches struggle with cue-level perception and reasoning due to two main challenges: 1) general-purpose modality encoders are pretrained to capture global structures and general semantics rather than fine-grained emotional cues, resulting in limited sensitivity to emotional signals; and 2) available datasets usually involve a trade-off between annotation quality and scale, which leads to insufficient supervision for emotional cues and ultimately limits cue-level reasoning. Moreover, existing evaluation metrics are inadequate for assessing cue-level reasoning performance. To address these challenges, we propose eXplainable Emotion GPT (XEmoGPT), a novel EMER framework capable of both perceiving and reasoning over emotional cues. It incorporates two specialized modules: the Video Emotional Cue Bridge (VECB) and the Audio Emotional Cue Bridge (AECB), which enhance the video and audio encoders through carefully designed tasks for fine-grained emotional cue perception. To further support cue-level reasoning, we construct a large-scale dataset, EmoCue, designed to teach XEmoGPT how to reason over multimodal emotional cues. In addition, we introduce EmoCue-360, an automated metric that extracts and matches emotional cues using semantic similarity, and release EmoCue-Eval, a benchmark of 400 expert-annotated samples covering diverse emotional scenarios. Experimental results show that XEmoGPT achieves strong performance in both emotional cue perception and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。