arXiv:2412.16694cs.CLcs.IR2024-12被引 1

构建《龙之家族》《权力的游戏》长文本问答数据集,提升叙事理解能力。

DragonVerseQA: Open-Domain Long-Form Context-Aware Question-Answering

  • 融合剧集摘要、影评、维基等多源数据,构建丰富上下文
  • 答案长度和情境复杂度显著高于SQuAD等主流数据集
  • 适合对话系统、剧情分析与情感计算研究者使用

本文提出一种新方法,构建面向奇幻宇宙《龙之家族》与《权力的游戏》电视系列的开放域长文本问答数据集DragonVerseQA。现有多数问答数据集聚焦于源自维基百科的简短事实型答案,缺乏深度与情境丰富性,难以支持复杂叙事理解。本研究整合来自HBO与粉丝维基网站的完整剧集摘要、IMDb与烂番茄等平台的用户评论,以及来自WikiData等仓库的结构化数据,形成多维度上下文数据集。经严格数据清洗与筛选后,确保高质量、无偏见的评论可用,为长文本回答提供支撑。该数据集能有效支持对话系统、叙事分析、情感分析、摘要生成与关系抽取等任务。与SQuAD 2.0、TriviaQA及Natural Questions等先进数据集对比表明,本数据集在上下文复杂度和答案长度方面具有显著优势。详尽评论进一步揭示观众情感与叙事解读,为特定领域问答树立新质量标准。本工作深化了对娱乐内容的理解,推动数字媒体环境中更智能、更具创造力的AI交互发展。

原文摘要 · Abstract (English)

This paper proposes a novel approach to develop an open-domain and long-form Over-The-Top (OTT) Question-Answering (QA) dataset, DragonVerseQA, specifically oriented to the fantasy universe of "House of the Dragon" and "Game Of Thrones" TV series. Most existing QA datasets focus on short, fact-based answers sourced almost solely from Wikipedia articles, devoid of depth and contextual richness for sophisticated narrative understanding. We curate a dataset that combines full episode summaries sourced from HBO and fandom wiki websites, user reviews from sources like IMDb and Rotten Tomatoes, and high-quality, open-domain, legally admissible sources, and structured data from repositories like WikiData into one dataset. The dataset provides a multi-dimensional context, reflecting complex character dynamics and plot developments from these varied sources. That means, on equal footing, only after heavy data preprocessing and filtering methods will meaningful, non-spam unbiased reviews be available in this enriched dataset. The comprehensive insights are given through the long-form answers generated from this enriched context. This is what makes this valuable dataset for improving conversational AI, narrative analysis, sentiment analysis, summarization techniques, and relation extraction. A comparative analysis with state-of-the-art QA datasets such as SQuAD 2.0, TriviaQA, and Natural Questions brings to light the unique advantages of our dataset in terms of contextual complexity and answer length. Detailed reviews add layers to audience sentiment and narrative interpretation, raising the bar for domain-specific QA with a new quality benchmark. Our work also allows a deeper understanding of entertainment-industry content and opens the door to more knowledgeable and creative AI-driven interactions within digital media environments.

问答系统长文本理解叙事分析多源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。