无需人工标注关系,让大模型自动理解多人对话中的说话人角色。
Contrastive Speaker-Aware Learning for Multi-party Dialogue Generation with LLMs
- 用说话人感知编码和对比学习,隐式建模对话连贯性与角色
- 在两个数据集上超越现有方法,提升流畅性、连贯性和多样性
- 适合需要高质量多人对话生成的场景,如客服或虚拟会议
多人对话生成因多说话人交互和交织的对话线而面临重大挑战。传统方法在捕捉这些复杂性方面表现不佳,尤其依赖人工标注的对话关系。本文提出说话人感知大模型(SA-LLM),利用预训练大语言模型和说话人感知对比学习策略应对这些挑战。SA-LLM采用说话人属性输入编码和对比学习目标,无需显式关系标注即可隐式学习上下文连贯性和说话人角色。在Ubuntu IRC和Movie Dialogues数据集上的大量实验表明,SA-LLM在自动评估和人工评估中均显著优于最先进基线,在流畅性、连贯性、信息量和回复多样性方面表现更优。消融实验和详细错误分析进一步验证了所提说话人感知训练方法的有效性,展示了其在不同说话人角色和上下文长度下的鲁棒性。结果强调了SA-LLM作为高质量多人对话生成的强大多样化且无需标注的解决方案的潜力。
原文摘要 · Abstract (English)
Multi-party dialogue generation presents significant challenges due to the complex interplay of multiple speakers and interwoven conversational threads. Traditional approaches often fall short in capturing these complexities, particularly when relying on manually annotated dialogue relations. This paper introduces Speaker-Attentive LLM (SA-LLM), a novel generative model that leverages pre-trained Large Language Models (LLMs) and a speaker-aware contrastive learning strategy to address these challenges. SA-LLM incorporates a speaker-attributed input encoding and a contrastive learning objective to implicitly learn contextual coherence and speaker roles without explicit relation annotations. Extensive experiments on the Ubuntu IRC and Movie Dialogues datasets demonstrate that SA-LLM significantly outperforms state-of-the-art baselines in automatic and human evaluations, achieving superior performance in fluency, coherence, informativeness, and response diversity. Ablation studies and detailed error analyses further validate the effectiveness of the proposed speaker-attentive training approach, highlighting its robustness across different speaker roles and context lengths. The results underscore the potential of SA-LLM as a powerful and annotation-free solution for high-quality multi-party dialogue generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。