多模态视频检索突破单模态限制,支持跨模态分段查询。
Enhanced Multimodal Video Retrieval System: Integrating Query Expansion and Cross-modal Temporal Event Retrieval

- 允许不同模态描述序列中不同场景,提升复杂时间上下文适应性。
- 用KDE-GMM算法自适应确定关键帧阈值,提高关键帧选取精度。
- 结合大模型扩写查询,适合需要精准定位的视频检索任务。
从视频中进行多媒体信息检索仍具挑战性。尽管近期系统已通过语义、物体和OCR查询实现多模态搜索,并能检索连续场景,但通常依赖单一查询模态处理整个序列,在复杂时间上下文中表现受限。为此,我们提出一种跨模态时间事件检索框架,使不同查询模态可分别描述序列中的不同场景。为自适应确定场景切换与画面切换的决策阈值,构建了核密度高斯混合阈值(KDE-GMM)算法,确保最优关键帧选择。提取的关键帧作为紧凑且高质量的视觉样本,保留各片段的语义核心,提升检索精度与效率。此外,系统引入大语言模型(LLM)对用户查询进行精炼与扩展,进一步增强整体检索性能。该系统在2025年胡志明市人工智能挑战赛中展现出优异效果。
原文摘要 · Abstract (English)
Multimedia information retrieval from videos remains a challenging problem. While recent systems have advanced multimodal search through semantic, object, and OCR queries - and can retrieve temporally consecutive scenes - they often rely on a single query modality for an entire sequence, limiting robustness in complex temporal contexts. To overcome this, we propose a cross-modal temporal event retrieval framework that enables different query modalities to describe distinct scenes within a sequence. To determine decision thresholds for scene transition and slide change adaptively, we build Kernel Density Gaussian Mixture Thresholding (KDE-GMM) algorithm, ensuring optimal keyframe selection. These extracted keyframes act as compact, high-quality visual exemplars that retain each segment's semantic essence, improving retrieval precision and efficiency. Additionally, the system incorporates a large language model (LLM) to refine and expand user queries, enhancing overall retrieval performance. The proposed system's effectiveness and robustness were demonstrated through its strong results in the Ho Chi Minh AI Challenge 2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。