用对比学习提升多人对话生成,兼顾说话风格与上下文关联
Advancing Multi-Party Dialogue Framework with Speaker-ware Contrastive Learning
- 通过两阶段自监督对比学习捕捉说话风格差异与话题转换
- 在多个基准上优于当前最优模型,且可增强大模型的多角色对话能力
- 适合研究多轮对话生成、个性化语音建模的研究者和工程师
多人对话常见于头脑风暴、谈判等协作场景,因其复杂性和发言者角色多样而极具挑战。现有方法多采用图神经网络建模对话上下文,虽能捕捉结构动态,但严重依赖标注图结构,且忽视个体说话风格。为此,我们提出CMR——一种基于对比学习的多人对话回复生成框架。该框架采用两阶段自监督对比学习:第一阶段捕捉个体间全局说话风格差异;第二阶段聚焦对话内比较,识别主题转移与语境相关事实。据我们所知,这是首个将对比学习应用于多人对话生成的方法。实验表明,CMR不仅显著优于当前最先进模型,还能有效提升大型预训练语言模型在多人对话中的表现,具备良好泛化能力。
原文摘要 · Abstract (English)
Multi-party dialogues, common in collaborative scenarios like brainstorming sessions and negotiations, pose significant challenges due to their complexity and diverse speaker roles. Current methods often use graph neural networks to model dialogue context, capturing structural dynamics but heavily relying on annotated graph structures and overlooking individual speaking styles. To address these challenges, we propose CMR, a Contrastive learning-based Multi-party dialogue Response generation framework. CMR employs a two-stage self-supervised contrastive learning framework. First, it captures global differences in speaking styles across individuals. Then, it focuses on intra-conversation comparisons to identify thematic transitions and contextually relevant facts. To the best of our knowledge, this is the first approach that applies contrastive learning in multi-party dialogue generation. Experimental results demonstrate that CMR not only significantly outperforms state-of-the-art models, but also generalizes well to large pre-trained language models, effectively enhancing their capability in handling multi-party conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。