arXiv:2507.12816cs.CVcs.AI2025-07

通过生成基础问题对提升视频理解,增强模型推理能力。

FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering

  • 基于视频描述自动生成基础问答对,补充场景上下文信息。
  • 在SUTD-TrafficQA上达到当前最优性能,显著提升泛化能力。
  • 适合需要强推理能力的视频问答任务研究者使用。

视频问答(VQA)是一项多模态任务,需理解视频内容以回答问题。现有方法主要依赖事件中心的问答对学习视频的时空特征,但这类标注缺乏物体类型、空间布局和描述性属性等关键细节,导致模型仅能学习碎片化的场景表示,限制了其泛化与高层推理能力。本文提出一种融合问题嵌入的基础问题生成方法(FIQ),通过从视频中提取描述生成问答对,丰富训练数据中的基础场景信息。生成的问答对使模型更好地理解视频主干上下文,从而提升泛化性和推理能力。此外,引入VQ-CAlign模块,将视觉特征与任务特定的问题嵌入对齐,保留领域特异性细节,增强下游任务适应性。在SUTD-TrafficQA数据集上的实验表明,FIQ优于现有基线方法,达到当前最佳性能。

原文摘要 · Abstract (English)

Video question answering (VQA) is a multimodal task that requires the interpretation of a video to answer a given question. Existing VQA methods primarily utilize question and answer (Q&A) pairs to learn the spatio-temporal characteristics of video content. However, these annotations are typically event-centric, which is not enough to capture the broader context of each video. The absence of essential details such as object types, spatial layouts, and descriptive attributes restricts the model to learning only a fragmented scene representation. This issue limits the model's capacity for generalization and higher-level reasoning. In this paper, we propose a fundamental question generation with the integration of question embeddings for video question answering (FIQ), a novel approach designed to strengthen the reasoning ability of the model by enhancing the fundamental understanding of videos. FIQ generates Q&A pairs based on descriptions extracted from videos, enriching the training data with fundamental scene information. Generated Q&A pairs enable the model to understand the primary context, leading to enhanced generalizability and reasoning ability. Furthermore, we incorporate a VQ-CAlign module that assists task-specific question embeddings with visual features, ensuring that essential domain-specific details are preserved to increase the adaptability of downstream tasks. Experiments on SUTD-TrafficQA demonstrate that our FIQ achieves state-of-the-art performance compared to existing baseline methods.

视频问答问题生成多模态推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。