构建视频驱动对话数据集与评估指标,提升活动类问答理解能力。
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
- 构建包含3000条对话的视频对话数据集,覆盖1000段复杂活动视频。
- 新评估指标与人工评分相关性显著高于现有方法。
- 适合研究多轮对话与视觉上下文理解的学者使用。
本文提出VDAct,一个面向事件驱动活动的视频接地对话数据集,以及针对该任务设计的会话级上下文评估指标VDEval。与现有数据集不同,VDAct包含更长、更复杂的视频序列,涵盖多种事件驱动活动,要求高级上下文理解以生成准确回复。数据集包含3,000条对话,超过30,000个问答对,源自1,000段具有多样活动场景的视频。由于活动场景广泛、问题类型丰富,该数据集具有显著挑战性。对主流视觉基础模型的实证研究表明,其在特定问题类型上仍存在明显局限。此外,VDEval通过融合对话会话历史和从补充知识图谱提取的视频内容摘要来评估单个回复,与人工评估的相关性显著高于仅依赖单轮对话上下文的现有指标。
原文摘要 · Abstract (English)
This paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for the task. Unlike existing datasets, VDAct includes longer and more complex video sequences that depict a variety of event-driven activities that require advanced contextual understanding for accurate response generation. The dataset comprises 3,000 dialogues with over 30,000 question-and-answer pairs, derived from 1,000 videos with diverse activity scenarios. VDAct displays a notably challenging characteristic due to its broad spectrum of activity scenarios and wide range of question types. Empirical studies on state-of-the-art vision foundation models highlight their limitations in addressing certain question types on our dataset. Furthermore, VDEval, which integrates dialogue session history and video content summaries extracted from our supplementary Knowledge Graphs to evaluate individual responses, demonstrates a significantly higher correlation with human assessments on the VDAct dataset than existing evaluation metrics that rely solely on the context of single dialogue turns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。