arXiv:2605.13803cs.CV2026-05

无需人工标注,通过双智能体自进化实现视频时序定位

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding

论文配图:EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
图 1 · 摘自论文原文
  • 用两个互演的智能体自动生成并优化查询-片段配对
  • 在2500段无标注视频上达到顶尖性能,超越有监督模型
  • 适合零样本视频理解、自监督学习研究者

视频时序定位(VTG)任务给定一段未剪辑视频和自然语言查询,需定位出最匹配查询的时间片段。现有方法依赖大量需人工标注的任务特定数据集,成本高昂。本文提出EvoGround框架,包含一个提议者与一个求解者两个耦合的自进化智能体,从原始视频中无须任何人工标注即可学习时序定位。提议者从原始视频中生成查询-时间段配对,求解者则学习进行定位,并将反馈信号回传以改进提议者。通过这一自我强化的强化学习循环,两个智能体从同一骨干网络初始化,在迭代中相互提升。在2500段未标注视频上训练后,EvoGround在多个VTG基准上表现匹配或超越完全监督模型,同时作为无标注细粒度视频描述生成器也达到前沿水平。

原文摘要 · Abstract (English)

Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task-specific datasets requiring costly manual annotation. We introduce EvoGround, a framework of two coupled self-evolving agents, a proposer and a solver, that learn temporal grounding from raw videos without any human-labeled data. The proposer generates query--moment pairs from raw videos, while the solver learns to ground them and feeds back signals that improve the proposer in return. Through this self-reinforcing reinforcement-learning loop, the two agents are initialized from the same backbone and mutually improve across iterations. Trained on 2.5K unlabeled videos, EvoGround matches or surpasses fully supervised models across multiple VTG benchmarks, while emerging as a state-of-the-art fine-grained video captioner without manual labels.

视频定位自进化自监督零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。