arXiv:2512.12929cs.CVcs.AI2025-12

让视频检索理解多事件时间关系,还能补全罕见概念。

MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation

  • 用连续片段相似度聚合,捕捉事件间时间连贯性。
  • 引入网络图像搜索扩展查询,提升对罕见概念的识别能力。
  • 适合需要跨事件、跨概念检索的智能视频系统开发者。

在线平台视频内容的快速增长,推动了对能够理解孤立视觉时刻及复杂事件时间结构的检索系统的需求。现有方法在建模多事件间的时间依赖关系,以及处理涉及未见或罕见视觉概念的查询方面表现不足。为此,我们提出由AIO_Trinh团队开发的MADTempo视频检索框架,将时间搜索与大规模视觉定位相结合。其时间搜索机制通过聚合连续视频片段的相似度得分,捕捉事件级连续性,实现多事件查询的连贯检索。同时,基于Google Image Search的回退模块,利用外部网络图像扩展查询表征,有效弥补预训练视觉嵌入的缺失,增强对分布外(OOD)查询的鲁棒性。二者协同提升了现代视频检索系统的时间推理与泛化能力,为大规模视频语料库中更语义化、自适应的检索铺平道路。

原文摘要 · Abstract (English)

The rapid expansion of video content across online platforms has accelerated the need for retrieval systems capable of understanding not only isolated visual moments but also the temporal structure of complex events. Existing approaches often fall short in modeling temporal dependencies across multiple events and in handling queries that reference unseen or rare visual concepts. To address these challenges, we introduce MADTempo, a video retrieval framework developed by our team, AIO_Trinh, that unifies temporal search with web-scale visual grounding. Our temporal search mechanism captures event-level continuity by aggregating similarity scores across sequential video segments, enabling coherent retrieval of multi-event queries. Complementarily, a Google Image Search-based fallback module expands query representations with external web imagery, effectively bridging gaps in pretrained visual embeddings and improving robustness against out-of-distribution (OOD) queries. Together, these components advance the temporal reasoning and generalization capabilities of modern video retrieval systems, paving the way for more semantically aware and adaptive retrieval across large-scale video corpora.

视频检索时间建模查询增强多事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。