用多智能体框架提升多模态共情回复的准确性与自然度
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation

- 分步解析情绪线索,构建结构化共情推理链
- 通过全局反思模块减少情感偏差,生成更符合情境的回应
- 适合研究情感计算与对话系统的开发者参考
多模态共情回复生成(MERG)旨在基于用户的多模态上下文生成情感共鸣的回应。现有方法通常采用从多模态上下文到最终回复的隐式单次生成范式,忽略了两个内在特性:(1) 人类对情绪线索的感知具有结构性而非直接映射,传统范式忽视了情绪认知的层次演进,导致情感判断失真;(2) 由于人类情感固有的复杂性与模糊性,该范式易产生显著情感偏见,最终影响共情效果。本文提出一种多智能体框架,通过结构化推理与反思优化提升共情能力。首先引入结构化共情推理-生成模块,显式分解为多模态感知、一致性感知情绪预测、实用策略规划和策略引导生成四个阶段,明确从多模态证据到回应实现的中间路径。此外,设计全局反思与精炼模块,由全局反思智能体逐步审计中间状态与生成结果,消除情感偏见与共情错误,并触发针对性重生成。整体闭环框架使模型在迭代中逐步提升情绪感知准确率并消除情感偏差。在IEMOCAP与MELD等多个基准上的实验表明,本模型在共情回复生成方面优于当前最先进方法。
原文摘要 · Abstract (English)
Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Existing approaches usually rely on an implicit one-pass generation paradigm from multimodal context to the final response, which overlooks two intrinsic characteristics of MERG: (1) Human perception of emotional cues is inherently structured rather than a direct mapping. The conventional paradigm neglects the hierarchical progression of emotion perception, leading to distorted emotional judgments. (2) Given the inherent complexity and ambiguity of human emotions, the conventional paradigm is prone to significant emotional biases, ultimately resulting in suboptimal empathy. In this paper, we propose a multi-agent framework for MERG, which enhances empathy through structured reasoning and reflective refinement. Specifically, we first introduce a structured empathetic reasoning-to-generation module that explicitly decomposes response generation via multimodal perception, consistency-aware emotion forecasting, pragmatic strategy planning, and strategy-guided response generation, providing a clearer intermediate path from multimodal evidence to response realization. Besides, we develop a global reflection and refinement module, in which a global reflection agent performs step-wise auditing over intermediate states and the generated response, eliminating existing emotional biases and empathy errors, and triggering targeted regeneration. Overall, such a closed-loop framework enables our model to gradually improve the accuracy of emotion perception and eliminate emotion biases during the iteration process. Experiments on several benchmarks, e.g., IEMOCAP and MELD, demonstrate that our model has superior empathic response generation capabilities compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。