构建首个混合印地语与英语的多模态对话语篇分析数据集
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
- 混合印地语与英语的对话数据,含音频与转录文本
- 标注九种语篇关系,覆盖多领域真实对话场景
- 揭示现有模型在跨语言混用场景下的显著性能下降
语篇分析是摘要、机器理解与情感识别等自然语言理解应用的重要任务。当前基于对话的语篇分析数据集多为单一领域的书面英文对话。本文介绍 CoMuMDR:首个用于对话语篇分析的混合印地语与英语的多模态多领域语料库。该语料库包含音视频与对应转写文本,标注了九种语篇关系。我们测试了多种主流基线模型,结果表明现有先进模型表现不佳,凸显多领域混合语言语料带来的挑战,提示需开发更适应此类真实场景的新模型。
原文摘要 · Abstract (English)
Discourse parsing is an important task useful for NLU applications such as summarization, machine comprehension, and emotion recognition. The current discourse parsing datasets based on conversations consists of written English dialogues restricted to a single domain. In this resource paper, we introduce CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations. The corpus (code-mixed in Hindi and English) has both audio and transcribed text and is annotated with nine discourse relations. We experiment with various SoTA baseline models; the poor performance of SoTA models highlights the challenges of multi-domain code-mixed corpus, pointing towards the need for developing better models for such realistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。