arXiv:2605.21973cs.CV2026-05中稿 · ICML被引 1

将视频时间定位转化为可验证的先识别后度量流程,提升定位准确性与稳定性。

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

论文配图:Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
图 1 · 摘自论文原文
  • 先预测事件候选段,再基于证据推理边界,分离识别与定位任务
  • 在多个基准上实现更优的定位准确率,且对不同模型骨架具有强泛化能力
  • 适合需要高可靠性和可解释性的视频理解场景,如智能监控与医疗分析

当前基于视频大模型的视频时间定位方法通常直接从非结构化的视觉标记流生成时间戳,易导致数值脆弱和边界不一致。为此,我们提出Foresee-to-Ground(F2G)框架,将时间定位重构为可验证的“识别-测量”问题。F2G融合预测性时序感知与证据驱动推理:学习敏感于边界的时序表示,构建全局候选事件段证据池,并将这些段落作为可引用证据单元提供给大语言模型,使边界预测与明确事件假设绑定。通过解耦事件识别与精确边界测量,F2G提升了定位稳定性并增强可验证性。大量实验表明,F2G在多个基准上持续提升定位精度,对不同视频大模型骨干网络具备良好迁移能力,并保持通用视频理解性能。

原文摘要 · Abstract (English)

Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle numerics and inconsistent boundaries. To address this, we propose Foresee-to-Ground (F2G), a framework that reformulates VTG as a verifiable Identify-then-Measure problem. F2G integrates Predictive Temporal Perception with Evidence-Driven Reasoning: it learns boundary-sensitive temporal representations to build a video-wide evidence pool of candidate event segments, and exposes these segments to the LLM as citable evidence units that bind boundary prediction to explicit event hypotheses. By decoupling event identification from precise boundary measurement, F2G stabilizes grounding and makes predictions verifiable. Extensive experiments demonstrate that F2G consistently improves grounding accuracy across diverse benchmarks, transfers robustly across different Video-LLM backbones, and preserves general video understanding capabilities.

视频定位大模型可解释性时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。