arXiv:2602.13313cs.CVcs.AI2026-02被引 1

用智能体协作实现无需训练的视频目标定位,效率更高。

Agentic Spatio-Temporal Grounding via Collaborative Reasoning

  • 构建空间与时间双智能体,通过对话式推理定位目标
  • 在多个基准上超越弱监督与零样本方法,接近全监督性能
  • 无需训练、无需标注,适合开放世界场景应用

时空视频定位(STVG)旨在给定文本查询后,从视频中检索出目标物体或人物的时空轨迹。现有方法通常在预测的时间区间内逐帧进行空间定位,导致计算冗余、依赖大量标注且泛化能力有限。弱监督方法虽降低标注成本,但仍受限于数据集级别的训练-拟合范式,性能较差。为此,本文提出面向开放世界与无训练场景的智能体时空定位框架(ASTG)。该框架由基于多模态大模型的空间推理智能体(SRA)和时间推理智能体(TRA)协同工作,遵循“提出-评估”范式,解耦时空推理过程,自动完成轨迹提取、验证与时间定位。借助专用视觉记忆与对话上下文,显著提升检索效率。在主流基准上的实验表明,该方法优于现有弱监督与零样本方法,性能可媲美部分全监督方法。

原文摘要 · Abstract (English)

Spatio-Temporal Video Grounding (STVG) aims to retrieve the spatio-temporal tube of a target object or person in a video given a text query. Most existing approaches perform frame-wise spatial localization within a predicted temporal span, resulting in redundant computation, heavy supervision requirements, and limited generalization. Weakly-supervised variants mitigate annotation costs but remain constrained by the dataset-level train-and-fit paradigm with an inferior performance. To address these challenges, we propose the Agentic Spatio-Temporal Grounder (ASTG) framework for the task of STVG towards an open-world and training-free scenario. Specifically, two specialized agents SRA (Spatial Reasoning Agent) and TRA (Temporal Reasoning Agent) constructed leveraging on modern Multimoal Large Language Models (MLLMs) work collaboratively to retrieve the target tube in an autonomous and self-guided manner. Following a propose-and-evaluation paradigm, ASTG duly decouples spatio-temporal reasoning and automates the tube extraction, verification and temporal localization processes. With a dedicate visual memory and dialogue context, the retrieval efficiency is significantly enhanced. Experiments on popular benchmarks demonstrate the superiority of the proposed approach where it outperforms existing weakly-supervised and zero-shot approaches by a margin and is comparable to some of the fully-supervised methods.

视频定位智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。