让对话摘要同时捕捉内容与情绪变化轨迹
Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition

- 从话题和人物双视角分解对话,分别生成摘要
- 结合自动识别的情绪信息,还原对话情感流变
- 适合做情感计算、人机交互等领域的研究者
现有文本摘要研究多聚焦单向信息(如新闻、报告),未考虑说话人之间的互动。而对话是多方来回交流构建意义的重要渠道。本文提出一种对话摘要框架,基于改进的分层链式智能体方法,利用多模态对话输入显式建模语义与情绪动态。通过两种方式分解对话:(1) 基于所有参与者发言的话题片段;(2) 针对每个参与者的发言片段。分别生成对应摘要,并融合自动推断的情绪信息。再将话题级与人物级摘要聚合为整体对话摘要,以捕捉语义内容与情绪演变过程。为超越内容准确性评估,引入情绪轨迹度量指标,衡量摘要保留情感流动的能力。在小语言模型上对多模态对话数据集的实验表明,该框架能生成兼具语义与情绪信息的摘要。进一步实验验证了方法有效性,并展示了语言模型在对话分析中的潜力。
原文摘要 · Abstract (English)
Existing text summarization research has focused much on monologic information (e.g., newspaper articles, reports) without accounting for the interaction between speakers or authors. In contrast, dialogues are a rich communication channel where multiple participants conduct back and forth exchanges to construct meaning. We propose a dialogue summarization framework that explicitly models both semantic and emotion dynamics using multimodal dialogue inputs, built on an adapted hierarchical Chain-of-Agents approach. We decompose dialogues from two perspectives: (1) topic segments based on the utterances of all participants, and (2) participant-specific utterance segments. These are used to generate corresponding summaries while incorporating automatically inferred emotions. Topic- and participant-level summaries are aggregated into a dialogue summary capturing semantic content and emotion trajectories. To evaluate beyond content accuracy, we introduce emotion trajectory metrics measuring how well summaries preserve emotional flow. Experiments with small language models on multimodal dialogue datasets show that our framework produces summaries with both semantic and emotion content. Further experiments on explicit emotion label availability highlight the efficacy of our proposed methodology and the opportunities in dialogue analysis using language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。