用秒级追踪与强化学习,高效精准定位长视频中的目标
Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

- 从帧级推理改为秒级跟踪,大幅降低计算量
- 在多个帧率下实现高精度定位,效率与准确率平衡出色
- 适合需要快速处理长视频的多模态应用
长视频中的时空定位需精确的时间定位和鲁棒的目标跟踪,以响应自然语言查询。尽管近期视觉语言模型(VLMs)具备强大推理能力,但直接对长序列进行逐帧推理计算开销大且不稳定。本文提出一种实用流程:将推理层级从帧级提升至秒级,通过跨秒平滑保持连续性的同时显著减少序列长度。为增强推理监督,利用先进的多模态模型生成链式思维风格轨迹用于时间定位与目标选择,并将生成的时空坐标替换为真实标注以避免噪声干扰。进一步采用基于 $t\_\mathrm{IoU}+mv\_\mathrm{IoU}$ 的验证器优化策略。在多个帧率设置下的实验表明,该方法在效率与定位质量间取得优异权衡。
原文摘要 · Abstract (English)
Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language models (VLMs) show strong reasoning ability, directly applying frame-by-frame inference to long sequences is computationally expensive and unstable. We propose a practical pipeline that shifts from frame-level to second-level tracking and performs cross-second smoothing to preserve continuity while reducing sequence length. To improve reasoning supervision, we synthesize chain-of-thought style trajectories using advanced multimodal models for temporal localization and target selection, and replace generated spatio-temporal coordinates with ground-truth annotations to avoid noisy supervision. We further optimize the policy with reinforcement learning using a verifier based on $t\_\mathrm{IoU}+mv\_\mathrm{IoU}$. Experiments across multiple FPS settings show that our method achieves a strong trade-off between efficiency and localization quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。