arXiv:2506.09953cs.CVcs.AI2025-06

构建视频对话数据集,要求模型结合视觉信息与外部知识回答问题。

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos

  • 基于2017段视频构建对话数据集,每段含多轮问答。
  • 共40954轮对话,需结合视频片段与外部知识作答。
  • 适合研究视觉对话与知识融合的学者使用。

在外部知识视觉问答(OK-VQA)任务中,模型需从图像中识别相关视觉信息,并结合外部知识准确回答问题。将该任务扩展至基于视频的对话场景时,对话模型不仅要在时间序列中识别关键视觉细节,还需回答那些所需信息并不直接存在于视觉内容中的问题。此外,后续对话必须考虑整体对话上下文。为此,我们构建了一个包含2,017段视频、5,986组人工标注对话、共计40,954轮交错对话回合的数据集。尽管对话上下文在特定视频片段中具有视觉基础,但问题仍需依赖未在视觉中呈现的外部知识。因此,模型不仅要定位相关视频部分,还需利用外部知识进行有效对话。我们还提供了在该数据集上评估的多个基线方法,并揭示了未来研究面临的挑战。数据集已公开:https://github.com/c-patsch/OKCV。

原文摘要 · Abstract (English)

In outside knowledge visual question answering (OK-VQA), the model must identify relevant visual information within an image and incorporate external knowledge to accurately respond to a question. Extending this task to a visually grounded dialogue setting based on videos, a conversational model must both recognize pertinent visual details over time and answer questions where the required information is not necessarily present in the visual information. Moreover, the context of the overall conversation must be considered for the subsequent dialogue. To explore this task, we introduce a dataset comprised of $2,017$ videos with $5,986$ human-annotated dialogues consisting of $40,954$ interleaved dialogue turns. While the dialogue context is visually grounded in specific video segments, the questions further require external knowledge that is not visually present. Thus, the model not only has to identify relevant video parts but also leverage external knowledge to converse within the dialogue. We further provide several baselines evaluated on our dataset and show future challenges associated with this task. The dataset is made publicly available here: https://github.com/c-patsch/OKCV.

视频对话知识融合数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。