基于电影剧本构建对话数据集,让聊天更持久真实
SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script
- 从电影剧本提取人物共享记忆,构建长期对话数据集
- 实验证明共享记忆使对话更吸引人且可持续
- 提出EPISODE框架,有效管理对话中的共享经历
两个人之间的共享记忆能增强彼此关系,并促进持续对话。本研究通过利用这些共享记忆,提升长时对话的吸引力。为此,我们从电影剧本中构建了一个名为SHARE的新长时对话数据集,其中包含人物角色信息与事件摘要,以及显性与隐性共享记忆。我们还提出了EPISODE框架,基于SHARE进行长时对话建模,利用人物间的共同经历。实验表明,共享记忆显著提升了对话的吸引力和可持续性,且EPISODE能有效管理共享记忆。数据集与代码已开源。
原文摘要 · Abstract (English)
Shared memories between two individuals strengthen their bond and are crucial for facilitating their ongoing conversations. This study aims to make long-term dialogue more engaging by leveraging these shared memories. To this end, we introduce a new long-term dialogue dataset named SHARE, constructed from movie scripts, which are a rich source of shared memories among various relationships. Our dialogue dataset contains the summaries of persona information and events of two individuals, as explicitly revealed in their conversation, along with implicitly extractable shared memories. We also introduce EPISODE, a long-term dialogue framework based on SHARE that utilizes shared experiences between individuals. Through experiments using SHARE, we demonstrate that shared memories between two individuals make long-term dialogues more engaging and sustainable, and that EPISODE effectively manages shared memories during dialogue. Our dataset and code are available at https://github.com/e1kim/SHARE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。