arXiv:2412.18364cs.CL2024-12被引 1

从对话中提取结构化三元组,提升社交智能体的可解释性。

Extracting triples from dialogues for conversational social agents

  • 设计五种模型从对话中抽取语义三元组
  • 单句三元组准确率达51.14%,元素级达69.32%
  • 首次公开对话三元组数据集,适合对话系统研究者

在人机协同的混合智能系统中,获取对话内容的显式理解对构建可控、透明的智能体至关重要。本文提出一系列自然语言理解模型,用于从社交对话中提取显式的符号三元组。三元组抽取此前主要基于维基百科文本进行知识库补全训练与测试,但社交对话具有独特性:对话双方通过多轮话语交替传递信息,包含陈述、提问、回答等复杂互动形式。对话中指代消解、省略、并列、隐含或明确的否定与确认等现象远比维基文本更突出。为此,我们首次发布用于训练与测试对话三元组抽取的数据集,并构建五种抽取模型进行评估。在单句层面,完整三元组最高精确率为51.14%,三元组元素级精确率为69.32%;然而,跨多轮对话的三元组抽取表现显著下降,表明从真实对话中提取知识仍具挑战性。

原文摘要 · Abstract (English)

Obtaining an explicit understanding of communication within a Hybrid Intelligence collaboration is essential to create controllable and transparent agents. In this paper, we describe a number of Natural Language Understanding models that extract explicit symbolic triples from social conversation. Triple extraction has mostly been developed and tested for Knowledge Base Completion using Wikipedia text and data for training and testing. However, social conversation is very different as a genre in which interlocutors exchange information in sequences of utterances that involve statements, questions, and answers. Phenomena such as co-reference, ellipsis, coordination, and implicit and explicit negation or confirmation are more prominent in conversation than in Wikipedia text. We therefore describe an attempt to fill this gap by releasing data sets for training and testing triple extraction from social conversation. We also created five triple extraction models and tested them in our evaluation data. The highest precision is 51.14 for complete triples and 69.32 for triple elements when tested on single utterances. However, scores for conversational triples that span multiple turns are much lower, showing that extracting knowledge from true conversational data is much more challenging.

对话理解三元组抽取知识提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。