统一视频事件搜索框架,高效跨源检索复杂事件。
U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025

- 融合三模块为统一架构,支持多类型查询一致处理。
- 提出轻量级关键帧提取方法,基于JPEG大小变化识别场景切换。
- 引入时序一致的文本描述生成,提升检索上下文感知能力。
从大规模视频数据集中检索事件面临时空与多模态信息复杂的挑战。本文提出U-CESE,为AI Challenge HCMC 2025设计的统一剪辑式事件搜索引擎,实现跨多样化视频源的多模态事件检索。在原有CESE基础上,将三个模块整合为单一协同框架,确保不同查询类型的处理与检索一致性。核心组件为统一剪辑算法,将多个独立剪辑流程合并为高效流水线。针对大规模数据,提出DAKE——一种无需训练的轻量级关键帧提取方法,利用JPEG文件大小变化检测显著场景转换。最后,引入ReCap,一种受循环神经网络启发的时序一致描述生成框架,可生成详尽且具备上下文感知能力的文本描述。实验表明,U-CESE在大规模多模态事件检索中表现出鲁棒、一致且高效的性能。
原文摘要 · Abstract (English)
Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CESE, our solution for the AI Challenge HCMC 2025, a Unified Clip-based Event Search Engine for multimodal event retrieval across diverse video sources. Building on CESE, U-CESE integrates its three modules into a single cohesive framework, ensuring consistent processing and retrieval across query types. A core component is the Unified Clipping Algorithm, which merges separate clipping algorithms into one efficient pipeline. To handle large-scale data, we propose DAKE, a lightweight, training-free keyframe extraction method using JPEG file size variations to identify significant scene changes. Finally, we introduce ReCap, a temporally consistent captioning framework inspired by Recurrent Neural Network, generating detailed and context-aware textual descriptions. Experiments show that U-CESE delivers robust, consistent, and efficient performance in large-scale multimodal event retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。