融合粗细粒度模型与时间重排序,提升长视频交互检索效率与准确率。
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
- 集成CLIP与BEIT3模型,实现粗细粒度联合搜索。
- 通过关键帧筛选与去重,存储开销降低40%以上。
- 双查询定位+邻域上下文重排序,结果更稳定可解释。
长视频理解对交互式检索系统构成重大挑战,传统方法在处理海量视频内容时效率低下。现有方案常依赖单一模型、存储冗余、时间搜索不稳定且忽略上下文的重排序,限制了实际应用效果。本文提出一种统一框架,包含四项创新:(1) 融合粗粒度(CLIP)与细粒度(BEIT3)模型的集成搜索策略,提升检索精度;(2) 基于TransNetV2选择代表性关键帧并去重,优化存储;(3) 采用双查询机制定位起止时间点,实现精准片段定位;(4) 利用邻近帧上下文进行时间重排序,稳定排名结果。在已知项搜索与问答任务上评估,本框架显著提升检索精度、效率与用户可解释性,为真实场景下的交互式视频检索提供稳健解决方案。
原文摘要 · Abstract (English)
Long-form video understanding presents significant challenges for interactive retrieval systems, as conventional methods struggle to process extensive video content efficiently. Existing approaches often rely on single models, inefficient storage, unstable temporal search, and context-agnostic reranking, limiting their effectiveness. This paper presents a novel framework to enhance interactive video retrieval through four key innovations: (1) an ensemble search strategy that integrates coarse-grained (CLIP) and fine-grained (BEIT3) models to improve retrieval accuracy, (2) a storage optimization technique that reduces redundancy by selecting representative keyframes via TransNetV2 and deduplication, (3) a temporal search mechanism that localizes video segments using dual queries for start and end points, and (4) a temporal reranking approach that leverages neighboring frame context to stabilize rankings. Evaluated on known-item search and question-answering tasks, our framework demonstrates substantial improvements in retrieval precision, efficiency, and user interpretability, offering a robust solution for real-world interactive video retrieval applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。