arXiv:2503.15887cs.CV2025-03被引 7

首个面向文档视频的问答数据集,助力模型理解图文音混合信息。

DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering

  • 构建文档视频问答任务与数据集,涵盖1454个视频、154k问题对。
  • 提出DV-LLaMA模型,在多模态融合和时序理解上显著超越现有方法。
  • 适合研究多模态理解、教育视频分析或大模型应用的学者使用。

远程工作和在线课程已成为知识传播的重要方式,催生了大量以文档为基础的教学视频。这类视频以密集文本图像和音频为特征,信息高度关联视觉内容,对多模态理解能力要求更高,但因数据稀缺和复杂性,研究仍不充分。本文首次提出文档视频问答(DocVideoQA)任务与数据集,包含23类共1454个视频,总时长约828小时,标注154,000个问答对,涵盖人工与GPT生成。数据集评估模型在理解、时间感知和模态融合方面的能力。我们建立基于开源多模态大模型的基线,并提出DV-LLaMA,通过多样化指令微调数据增强单模态特征提取,结合对比学习强化跨模态整合。经微调后,大语言模型具备音视频联合理解能力,显著提升文档视频理解性能。在DocVideoQA上的大量实验表明,该方法优于现有模型。代码与数据集将公开,推动后续研究。

原文摘要 · Abstract (English)

Remote work and online courses have become important methods of knowledge dissemination, leading to a large number of document-based instructional videos. Unlike traditional video datasets, these videos mainly feature rich-text images and audio that are densely packed with information closely tied to the visual content, requiring advanced multimodal understanding capabilities. However, this domain remains underexplored due to dataset availability and its inherent complexity. In this paper, we introduce the DocVideoQA task and dataset for the first time, comprising 1454 videos across 23 categories with a total duration of about 828 hours. The dataset is annotated with 154k question-answer pairs generated manually and via GPT, assessing models' comprehension, temporal awareness, and modality integration capabilities. Initially, we establish a baseline using open-source MLLMs. Recognizing the challenges in modality comprehension for document-centric videos, we present DV-LLaMA, a robust video MLLM baseline. Our method enhances unimodal feature extraction with diverse instruction-tuning data and employs contrastive learning to strengthen modality integration. Through fine-tuning, the LLM is equipped with audio-visual capabilities, leading to significant improvements in document-centric video understanding. Extensive testing on the DocVideoQA dataset shows that DV-LLaMA significantly outperforms existing models. We'll release the code and dataset to facilitate future research.

文档视频多模态问答系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。