用多帧信息提升视频检索,让模型理解动作而非单图物体
Multimodal Contextualized Support for Enhancing Video Retrieval System
- 融合多帧视觉与上下文信息,构建视频级语义表示
- 相比单帧检索,显著提升对连续动作的识别准确率
- 适合需要理解视频事件的场景,如智能安防、内容推荐
当前视频检索系统(尤其竞赛中)主要依赖单个关键帧或图像查询,但实际查询常涉及一系列帧中的动作或事件。仅分析单帧导致信息不足,模型难以捕捉高层次抽象语义,通常仅描述帧内物体,缺乏深层理解。本文提出一种新系统,整合最新方法,设计新颖流水线:从视频片段中提取多模态数据,结合多帧信息,使模型能抽象出视频隐含意义,聚焦于可推断的事件整体,而非单一图像中的物体检测。
原文摘要 · Abstract (English)
Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encoding an entire clip or video segment. However, queries often describe an action or event over a series of frames, not a specific image. This results in insufficient information when analyzing a single frame, leading to less accurate query results. Moreover, extracting embeddings solely from images (keyframes) does not provide enough information for models to encode higher-level, more abstract insights inferred from the video. These models tend to only describe the objects present in the frame, lacking a deeper understanding. In this work, we propose a system that integrates the latest methodologies, introducing a novel pipeline that extracts multimodal data, and incorporate information from multiple frames within a video, enabling the model to abstract higher-level information that captures latent meanings, focusing on what can be inferred from the video clip, rather than just focusing on object detection in one single image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。