让AI同时懂情绪和通用视觉语言,不丢掉原有能力。
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
- 用专家混合机制动态分配处理任务,兼顾情绪与通用理解。
- 在4万+视频数据上训练,情绪识别准确率领先现有模型。
- 适合做情感分析、多模态交互的开发者和研究者使用。
视频中的精准情绪理解需融合视觉、文本、音频及上下文线索。尽管近期大型多模态模型(LMMs)在通用视觉-语言(VL)任务中表现优异,但在情绪相关任务中常出现性能下降,微调时易产生灾难性遗忘。为此,我们提出Emotion-Qwen,一个统一的多模态框架,可同时实现鲁棒的情绪理解与通用VL推理能力。该框架引入基于专家混合(MoE)架构的新型混合压缩器,动态路由输入,平衡情绪特异性处理与通用多模态推理。我们设计了三阶段预训练流程,利用大规模通用与情绪聚焦数据集,增强多模态表征鲁棒性与模型适应性。此外,我们构建了视频情绪推理(VER)数据集,包含超过4万段双语视频片段,附带详尽的情境感知情绪标注,显著推动细粒度情绪推理研究。大量实验表明,Emotion-Qwen在多个情绪识别与推理基准上达到顶尖水平,同时在通用VL任务中保持高度竞争力。
原文摘要 · Abstract (English)
Accurate emotion understanding in videos necessitates effectively recognizing and interpreting emotional states by integrating visual, textual, auditory, and contextual cues. Although recent Large Multimodal Models (LMMs) have exhibited significant progress in general vision-language (VL) tasks, their performance often deteriorates in emotion-specific scenarios, exhibiting catastrophic forgetting when fine-tuned on emotion-centric tasks. To overcome these limitations, we propose Emotion-Qwen, a unified multimodal framework designed to simultaneously enable robust emotion understanding and preserve general VL reasoning capabilities. Emotion-Qwen introduces a novel Hybrid Compressor based on a Mixture-of-Experts (MoE) architecture, dynamically routing inputs to optimally balance emotion-specific processing and general multimodal reasoning. We further propose a carefully structured three-stage pre-training pipeline, leveraging extensive general and emotion-focused datasets to strengthen multimodal representation robustness and model adaptability. Additionally, we develop the Video Emotion Reasoning (VER) dataset, a large-scale bilingual resource containing over 40K video clips annotated with detailed context-aware emotional descriptions, significantly facilitating research on fine-grained emotional reasoning. Extensive experiments confirm that Emotion-Qwen achieves state-of-the-art performance across multiple emotion recognition and reasoning benchmarks, while maintaining highly competitive results in general VL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。