arXiv:2412.19304cs.CV2024-12被引 4

提出T-Former,让视频问答更精准定位关键时间片段。

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries

  • 用问题引导时间建模,连接视觉感知与大模型推理
  • 在多个视频问答数据集上表现优于现有方法
  • 适合需要精准时序理解的视频分析任务

视频问答(Video QA)是一项挑战性任务,要求模型理解完整视频,根据问题上下文识别最相关的信息,并准确推理出答案。近年来,多模态大语言模型(MLLMs)凭借其出色的常识推理能力推动了该领域的发展,主要依赖于视觉数据与语言空间的有效对齐。然而,视频问答还需解决时空对齐难题,以从不同帧中提取与问题相关的内容。本文研究多种时间建模方法,结合MLLMs,实现基于问题引导的时间建模。我们提出T-Former,一种新颖的时间建模方法,构建帧级视觉感知与大语言模型推理能力之间的问题引导时间桥梁。在多个视频问答基准上的评估表明,T-Former在性能上可与现有先进方法竞争,并与当前视频问答技术进展保持一致。

原文摘要 · Abstract (English)

Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on contextual cues from a given question, and reason accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have transformed video QA by leveraging their exceptional commonsense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an additional space-time alignment poses a considerable challenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA benchmarks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with recent advancements in video QA.

视频问答时序建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。