arXiv:2505.01583cs.CVcs.AI2025-05被引 11

通过因果推理与细粒度分段提升视频理解能力

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

  • 用掩码事件预测重建缺失事件并生成因果解释
  • 在100万实例数据集上实现精准时间定位与事件分割
  • 适合需要精细动作分析的视频理解任务

视觉语言模型在理解因果事件关系和实现视频细粒度时间定位方面仍面临挑战。现有方法或压缩视频标记降低时间分辨率,或将视频视为无分割流,模糊了事件边界并限制了因果依赖建模。我们提出TEMPURA(Temporal Event Masked Prediction and Understanding for Reasoning in Action),一种两阶段训练框架,以增强视频时间理解。首先,通过掩码事件预测推理重建缺失事件,并基于密集事件标注生成逐步因果解释,借鉴高效补全技术;其次,学习执行视频分割与密集描述,将视频分解为非重叠事件并提供时间对齐的详细描述。我们在自建的大规模数据集VER上训练,该数据集包含100万训练样本和50万条具有时间对齐事件描述与结构化推理步骤的视频。在时间定位与亮点检测基准测试中,TEMPURA显著优于强基线模型,验证了将因果推理与细粒度时间分割结合可有效提升视频理解能力。

原文摘要 · Abstract (English)

Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. Existing methods either compress video tokens to reduce temporal resolution, or treat videos as unsegmented streams, which obscures fine-grained event boundaries and limits the modeling of causal dependencies. We propose TEMPURA (Temporal Event Masked Prediction and Understanding for Reasoning in Action), a two-stage training framework that enhances video temporal understanding. TEMPURA first applies masked event prediction reasoning to reconstruct missing events and generate step-by-step causal explanations from dense event annotations, drawing inspiration from effective infilling techniques. TEMPURA then learns to perform video segmentation and dense captioning to decompose videos into non-overlapping events with detailed, timestamp-aligned descriptions. We train TEMPURA on VER, a large-scale dataset curated by us that comprises 1M training instances and 500K videos with temporally aligned event descriptions and structured reasoning steps. Experiments on temporal grounding and highlight detection benchmarks demonstrate that TEMPURA outperforms strong baseline models, confirming that integrating causal reasoning with fine-grained temporal segmentation leads to improved video understanding.

视频理解因果推理事件分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。