arXiv:2410.20766cs.CLcs.AI2024-10被引 15

提出静态动态注意力框架,提升多轮对话连贯性

A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation

  • 设计静态与动态双重注意力机制建模对话历史
  • 在Ubuntu和Opensubtitles数据集上优于基线模型
  • 适合需要长程上下文理解的对话系统研究者

近年来,开放域对话系统受到学术界和产业界的广泛关注。其目标是模仿人类进行自然对话。单轮对话生成的研究已取得显著进展,但多轮对话因具有连贯性和上下文依赖性,难以通过单轮理解实现。因此,在开放域多轮对话生成中,建模对话历史的语义至关重要。已有研究验证了层次化循环编码器-解码器框架的有效性,但基于RNN的层级编码仍存在梯度消失问题。为此,本文提出一种基于静态与动态注意力的对话历史建模方法,用于生成开放域多轮对话响应。在Ubuntu和Opensubtitles数据集上的实验结果表明,该方法在自动评估与人工评估指标上均表现优异,且在多种实验设置下均有效。同时,实证验证了静态与动态注意力结合的协同优势。

原文摘要 · Abstract (English)

Recently, research on open domain dialogue systems have attracted extensive interests of academic and industrial researchers. The goal of an open domain dialogue system is to imitate humans in conversations. Previous works on single turn conversation generation have greatly promoted the research of open domain dialogue systems. However, understanding multiple single turn conversations is not equal to the understanding of multi turn dialogue due to the coherent and context dependent properties of human dialogue. Therefore, in open domain multi turn dialogue generation, it is essential to modeling the contextual semantics of the dialogue history, rather than only according to the last utterance. Previous research had verified the effectiveness of the hierarchical recurrent encoder-decoder framework on open domain multi turn dialogue generation. However, using RNN-based model to hierarchically encoding the utterances to obtain the representation of dialogue history still face the problem of a vanishing gradient. To address this issue, in this paper, we proposed a static and dynamic attention-based approach to model the dialogue history and then generate open domain multi turn dialogue responses. Experimental results on Ubuntu and Opensubtitles datasets verify the effectiveness of the proposed static and dynamic attention-based approach on automatic and human evaluation metrics in various experimental settings. Meanwhile, we also empirically verify the performance of combining the static and dynamic attentions on open domain multi turn dialogue generation.

多轮对话注意力机制对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。