提出新基准与框架,精准定位长视频中多个符合复杂条件的事件。
Conditional Multi-Event Temporal Grounding in Long-Form Video

- 将任务转化为结构化搜索与聚合,无需训练
- 在600段平均33.8分钟视频上实现6.1%的F1提升
- 适合需要高精度多事件定位的研究者与开发者
多模态大语言模型在视频时间定位上进展迅速,但真实场景常需定位满足复合时空条件的多个事件。现有基准存在局限:仅定位单个时刻、忽略时间条件或割裂定位与计数任务。我们构建了CoMET-Bench,包含600段平均33.8分钟长视频中的2789个查询,覆盖五个真实场景,每个查询由4个时间条件、3个空间条件构成,并设有专门的负向查询子集。提出统一评估协议,联合衡量计数、定位与负向查询识别,引入新的拒绝-准确率(Rejection-F1)以防止“始终为空”模型的作弊行为。对多种MLLM、基于代理及专用定位方法的评测显示,现有方法仍远未解决此任务。基于此,我们提出CoMET-Agent——一种无需训练的代理式框架,将任务重构为结构化搜索与聚合,纯靠推理使[email protected]提升6.1%。失败分析揭示三个开放方向:细粒度实体追踪、位置均匀检索与因果事件配对。
原文摘要 · Abstract (English)
Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositional temporal and spatial conditions. Existing benchmarks fall short: they localize only a single moment per query, count without temporal conditions, or treat grounding and counting as disjoint tasks. We introduce CoMET-Bench for Conditional Multi-Event Temporal Grounding in long-form video, comprising 2789 queries over 600 videos averaging 33.8 minutes across five real-world domains, with each query composed from 4 temporal conditions, 3 spatial conditions, and a dedicated negative-query subset. We further propose a unified evaluation protocol jointly measuring counting, grounding, and negative-query recognition, including a new Rejection-F1 metric that prevents trivial gaming by lazy "always-empty" models. Benchmarking a broad suite of MLLMs, agent-based, and grounding-specialized methods reveals that existing approaches remain far from solving this task. Building on these findings, we propose CoMET-Agent, a training-free agentic framework that reformulates the task as structured search-and-aggregate, improving [email protected] by 6.1% over GPT-5 purely through structural reasoning. Failure analysis further surfaces three open directions: fine-grained entity tracking, position-uniform retrieval, and causal event pairing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。