arXiv:2503.12855cs.CVcs.AI2025-03CVPR被引 6

让AI学会从长视频中找关键时间片段并推理,提升复杂问答能力

VITED: Video Temporal Evidence Distillation

  • 通过搜索最优证据区间,自动构建视频问答的推理链
  • 在长视频问答任务上超越现有方法,显著提升多步推理准确率
  • 适合需要精准时间定位与复杂逻辑推理的研究者和开发者

我们研究基于证据链推理的复杂视频问答——从视频多个相关片段中识别出时间跨度序列,并提取其中的视觉证据。现有模型因固定帧数均匀采样,难以捕捉非均匀分布的关键证据,且缺乏在完整视频上下文中精确定位证据的能力,导致多步推理表现不佳。我们提出框架,在现有VideoQA数据集中注入证据推理链,通过搜索支持答案的最佳兴趣区间,最大化问题回答概率。训练模型VITED直接生成这些证据链,使其既能定位证据窗口,也能在长视频内容中完成跨窗口多步推理。我们在一系列长视频问答基准上验证了该方法的价值,结果表明其性能优于缺乏证据推理能力的先进模型。

原文摘要 · Abstract (English)

We investigate complex video question answering via chain-of-evidence reasoning -- identifying sequences of temporal spans from multiple relevant parts of the video, together with visual evidence within them. Existing models struggle with multi-step reasoning as they uniformly sample a fixed number of frames, which can miss critical evidence distributed nonuniformly throughout the video. Moreover, they lack the ability to temporally localize such evidence in the broader context of the full video, which is required for answering complex questions. We propose a framework to enhance existing VideoQA datasets with evidence reasoning chains, automatically constructed by searching for optimal intervals of interest in the video with supporting evidence, that maximizes the likelihood of answering a given question. We train our model (VITED) to generate these evidence chains directly, enabling it to both localize evidence windows as well as perform multi-step reasoning across them in long-form video content. We show the value of our evidence-distilled models on a suite of long video QA benchmarks where we outperform state-of-the-art approaches that lack evidence reasoning capabilities.

视频问答推理链时间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。