通过因果推理精准定位教学视频中的回答片段,提升长视频问答准确率。
CACR:Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

- 基于视觉语言预训练生成候选片段,再用因果推理精炼答案。
- 在6个数据集上达到最高mIoU,显著优于现有方法。
- 适合需要长视频精准检索的场景,如智能教育系统。
指令性视频中的时间答案定位(TAGV)任务旨在精确定位回应自然语言查询的视频片段,对直接视频问答检索至关重要。该任务因需理解语义复杂的提问,且未剪辑视频与目标片段间存在显著长度差异而极具挑战。现有方法常受无关内容干扰或视觉推理能力不足影响。为此,我们提出候选感知因果推理(CACR)框架:首先利用基于视觉语言预训练的候选选择(VBCS)算法高效生成K个候选片段;随后通过增强排斥奖励机制的时序逻辑推理模块,并采用分组相对策略优化(GRPO)进行优化,实现鲁棒推理。在六个基准数据集上的大量实验表明,该方法在平均交并比(mIoU)上达到当前最优性能,为长视频中基于推理的检索提供了新视角。
原文摘要 · Abstract (English)
The task of temporal answer grounding in instructional video (TAGV), which aims to locate precise video segments that respond to natural language queries, is increasingly important for direct video answer retrieval. This task remains challenging due to the need to comprehend semantically complex questions and to address the significant length mismatch between untrimmed videos and short target moments. Existing methods often suffer from sensitivity to irrelevant content or insufficient visual reasoning capabilities. To tackle these limitations, we propose a Candidate-Aware Causal Reasoning (CACR) framework. Our approach first employs a Visual-Language Pre-training based Candidate Selection (VBCS) algorithm to efficiently generate K candidate segments, then applies a temporal logic reasoning module enhanced by a rejection reward mechanism and optimized via Group Relative Policy Optimization (GRPO) for robust inference. Extensive experiments on six benchmarks demonstrate that our method achieves state-of-the-art performance in terms of mean Intersection-over-Union (mIoU), providing a new perspective for reasoning-based retrieval in long videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。