首个多角色多模态对话话语解析数据集,助力对话结构理解
DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset

- 基于美剧构建多角色多模态对话数据集
- 含495段对话6374条语句,时长9.1小时
- 适合研究多模态对话理解的学者使用
多角色对话话语解析旨在识别对话中话语之间的依赖结构和关系类型。以往研究多局限于文本模态或双人对话,难以应对多模态与多角色场景。本文基于美国电视剧构建首个公开的英文多模态多角色对话话语解析数据集DraDDP,包含495个对话片段、6,374条语句及9.1小时并行视频内容,涵盖丰富的多角色互动情境。同时,我们在DraDDP上建立全面基准,评估该任务并深入分析不同模态的影响。实验表明,多模态信息在捕捉对话结构与关系类型方面具有重要价值。我们将公开数据集、标注指南与代码,推动多模态对话理解研究发展。
原文摘要 · Abstract (English)
Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations. Previous studies are mostly limited to textual modality or two-party dialogue, failing to meet the multimodal and multi-party settings. In this paper, we construct the first publicly available English multimodal dataset DraDDP for multi-party dialogue discourse parsing, based on American TV dramas. DraDDP contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, covering rich multi-party interaction scenarios. Moreover, we establish comprehensive benchmarks by evaluating this task on DraDDP and conducting in-depth analysis on the impact of different modalities. Experimental results demonstrate the value of multimodal information in capturing dialogue structures and relation types. We will publicly release the dataset, annotation guidelines, and code to promote future research in multimodal dialogue understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。