用强化学习动态分工,让多模态模型更懂复杂情绪。
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

- 通过强化学习实现情绪识别代理的动态分工
- 在多个基准上情绪推理能力显著提升
- 适合研究多模态情感分析与智能代理系统的人
多模态大语言模型(MLLM)在多模态情感识别(MER)任务中表现卓越,实现了复杂情绪理解与视频内容解析能力。然而,现有方法通常使用固定提示感知情绪,忽视了多模态输入中情感源的动态性与复杂性。为此,我们提出一种基于强化学习的动态代理专业化框架(EmoAgent-R1),以优化MLLM的情绪识别、推理与泛化能力。首先采用冷启动策略,通过合成答案引导的思维链数据和代理路由数据训练,赋予模型初步情绪识别与代理调度能力;随后引入强化学习,在两步代理工作流中实现代理选择与专业化。为有效训练,我们设计了新型渐进式组相对策略优化(P-GRPO),结合组内相对优势与受PMI启发的渐进式词元级调制,将稀疏奖励转化为细粒度学习信号,缓解了传统GRPO中粗粒度统一信用分配的问题。在多个MER基准上的实验表明,EmoAgent-R1在情绪推理性能与优化稳定性方面均优于现有方法。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。