通过自动生成基础问题提升视频理解,增强模型推理能力
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
- 从视频中提取描述性信息生成基础问答对,补充场景级属性
- 在SUTD-TrafficQA上达到当前最佳性能,显著提升推理效果
- 适合需要强场景理解的视频问答研究者使用
传统视频问答(VQA)方法主要依赖事件为中心的问答对来学习视频的时空动态。然而,现有标注多聚焦于事件,限制了模型对场景完整上下文的理解。缺乏物体类别、空间布局和视觉属性等基础信息,导致模型难以形成全面环境认知,影响泛化与推理能力。本文提出嵌入融合式基础问题生成框架FIQ,直接从视频中提取描述性信息生成问答对,丰富数据集中的场景级核心属性。这些生成的问答对帮助模型建立更完整的视频理解,从而提升泛化性与推理表现。此外,我们设计了VQ-CAlign模块,将任务特定的问题嵌入与视觉特征对齐,保留关键上下文线索,增强下游任务适应性。在SUTD-TrafficQA数据集上的实验表明,FIQ达到当前最优性能,超越已有基线方法。
原文摘要 · Abstract (English)
Conventional VQA approaches primarily rely on question-answer (Q&A) pairs to learn the spatio-temporal dynamics of video content. However, most existing annotations are event-centric, which restricts the model's ability to capture the comprehensive context of a scene. The lack of fundamental information such as object categories, spatial configurations, and descriptive visual attributes prevents the model from forming a complete understanding of the environment, ultimately limiting its generalization and reasoning capability. In this paper, we introduce Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach (FIQ), a framework designed to enhance the reasoning capability of VQA models by improving their foundational comprehension of video content. FIQ generates Q&A pairs from descriptive information extracted directly from videos, thereby enriching the dataset with core scene-level attributes. These generated pairs help the model develop a more holistic understanding of the video, leading to improved generalizability and reasoning performance. In addition, we propose a VQ-CAlign module that aligns task-specific question embeddings with corresponding visual features, preserving essential contextual cues and enhancing adaptability to downstream tasks. Experimental results on the SUTD-TrafficQA dataset demonstrate that FIQ achieves state-of-the-art performance, surpassing existing baseline approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。