arXiv:2503.06820cs.CVcs.AI2025-03被引 2

构建细粒度视频问答数据集并提出新模型,提升对动作与关系的精准理解。

Towards Fine-Grained Video Question Answering

  • 引入包含场景图和时间区间标注的MOMA-QA数据集,聚焦时空细节。
  • 提出的SGVLM模型在多个数据集上达到新基准,尤其擅长定位与关系推理。
  • 适合研究视频理解、细粒度问答及多模态模型的开发者参考。

在快速发展的视频理解领域,视频问答(VideoQA)仍是核心任务。然而现有数据集在时间与空间粒度上存在不足,限制了现有方法的能力。本文提出多对象多角色问答数据集MOMA-QA,通过强调时间定位、空间关系推理和以实体为中心的查询,弥补上述缺陷。该数据集配备真实场景图和时间区间标注,适用于细粒度视频理解模型的开发。此外,我们提出一种新型视频-语言模型SGVLM,融合场景图预测器、高效帧检索器与预训练大语言模型,实现精准的时间定位与细粒度关系理解。在MOMA-QA及其他公开数据集上的评估表明,该模型性能优越,为视频问答设定了新基准。

原文摘要 · Abstract (English)

In the rapidly evolving domain of video understanding, Video Question Answering (VideoQA) remains a focal point. However, existing datasets exhibit gaps in temporal and spatial granularity, which consequently limits the capabilities of existing VideoQA methods. This paper introduces the Multi-Object Multi-Actor Question Answering (MOMA-QA) dataset, which is designed to address these shortcomings by emphasizing temporal localization, spatial relationship reasoning, and entity-centric queries. With ground truth scene graphs and temporal interval annotations, MOMA-QA is ideal for developing models for fine-grained video understanding. Furthermore, we present a novel video-language model, SGVLM, which incorporates a scene graph predictor, an efficient frame retriever, and a pre-trained large language model for temporal localization and fine-grained relationship understanding. Evaluations on MOMA-QA and other public datasets demonstrate the superior performance of our model, setting new benchmarks for VideoQA.

视频问答细粒度理解场景图多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。