arXiv:2409.07672cs.CLcs.AI2024-09被引 2

通过重写对话修复指代和省略,提升无监督话题分割效果

An Unsupervised Dialogue Topic Segmentation Model Based on Utterance Rewriting

  • 用话语重写技术恢复对话中的指代和省略信息
  • 在DialSeg711上绝对误差降低6%,在Doc2Dial上提升3%
  • 适合研究对话理解与无监督学习的学者参考

对话话题分割在各类对话建模任务中至关重要。现有无监督方法通过相邻话语匹配和伪分割学习话题感知的语篇表示,以挖掘未标注对话关系中的有用线索。然而,在多轮对话中,话语常存在指代或省略,直接使用原始话语进行表征学习可能影响邻近话语匹配中的语义相似性计算。为此,本文提出一种新颖的无监督对话话题分割方法,结合话语重写(UR)技术与无监督学习算法,通过重写对话以恢复指代和缺失词汇,从而更有效地利用未标注对话中的有用线索。相比现有模型,提出的话语重写话题分割模型(UR-DTS)显著提升话题分割准确率。在DialSeg711上,绝对误差分数提升约6%,达到11.42%;在WD指标上达12.97%。在Doc2Dial上,绝对误差分数和WD分别提升约3%和2%,分别达到35.17%和38.49%,达到当前最优性能。结果表明该模型能有效捕捉对话话题的细微差别,也揭示了利用未标注对话的潜力与挑战。

原文摘要 · Abstract (English)

Dialogue topic segmentation plays a crucial role in various types of dialogue modeling tasks. The state-of-the-art unsupervised DTS methods learn topic-aware discourse representations from conversation data through adjacent discourse matching and pseudo segmentation to further mine useful clues in unlabeled conversational relations. However, in multi-round dialogs, discourses often have co-references or omissions, leading to the fact that direct use of these discourses for representation learning may negatively affect the semantic similarity computation in the neighboring discourse matching task. In order to fully utilize the useful cues in conversational relations, this study proposes a novel unsupervised dialog topic segmentation method that combines the Utterance Rewriting (UR) technique with an unsupervised learning algorithm to efficiently utilize the useful cues in unlabeled dialogs by rewriting the dialogs in order to recover the co-referents and omitted words. Compared with existing unsupervised models, the proposed Discourse Rewriting Topic Segmentation Model (UR-DTS) significantly improves the accuracy of topic segmentation. The main finding is that the performance on DialSeg711 improves by about 6% in terms of absolute error score and WD, achieving 11.42% in terms of absolute error score and 12.97% in terms of WD. on Doc2Dial the absolute error score and WD improves by about 3% and 2%, respectively, resulting in SOTA reaching 35.17% in terms of absolute error score and 38.49% in terms of WD. This shows that the model is very effective in capturing the nuances of conversational topics, as well as the usefulness and challenges of utilizing unlabeled conversations.

对话分割无监督学习话语重写语篇分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。