构建对话语料库,研究日常对话中话题识别与转换机制。
"Wait, did you mean the doctor?": Collecting a Dialogue Corpus for Topical Analysis
- 用自研消息工具采集真实对话数据
- 聚焦多话题交替的长对话场景
- 为话题分析提供高质量语料支持
对话是人类行为的核心,识别当前话题对参与交流至关重要。然而,现有文献对日常对话中的话题组织方式及话题识别机制的研究仍很有限。话题分析需要包含多个话题和类型转换的长对话,这类数据的收集与标注十分困难。本文提出一项对话语料收集实验,旨在构建适合话题分析的语料库。我们将使用自研的消息工具进行数据采集,以获取具备丰富话题演变特征的真实对话数据。
原文摘要 · Abstract (English)
Dialogue is at the core of human behaviour and being able to identify the topic at hand is crucial to take part in conversation. Yet, there are few accounts of the topical organisation in casual dialogue and of how people recognise the current topic in the literature. Moreover, analysing topics in dialogue requires conversations long enough to contain several topics and types of topic shifts. Such data is complicated to collect and annotate. In this paper we present a dialogue collection experiment which aims to build a corpus suitable for topical analysis. We will carry out the collection with a messaging tool we developed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。