arXiv:2512.12935cs.CVcs.AI2025-12中稿 · AAAI

统一交互式视频片段检索,自动处理模糊查询并精准定位连贯事件序列。

Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion

  • 分层双嵌入+重排序,融合多模态特征提升检索覆盖与精度。
  • 时序感知评分机制通过指数衰减惩罚长间隔,生成连贯事件序列。
  • 智能代理自动拆解查询,动态融合视觉/文字/语音模态,无需人工选择。

视频内容的爆炸式增长催生了高效多模态片段检索系统的迫切需求。现有方法面临三大挑战:(1) 固定权重融合策略在跨模态噪声和模糊查询下表现不佳;(2) 时序建模难以捕捉连贯事件序列,且对不合理时间间隔惩罚不足;(3) 系统需手动选择模态,降低可用性。本文提出一种统一的多模态片段检索系统,包含三项创新:(1) 采用级联双嵌入流水线,结合BEIT-3与SigLIP实现广域检索,并通过基于BLIP-2的重排序平衡召回率与精确率;(2) 引入时序感知评分机制,利用束搜索施加指数衰减惩罚大时间间隔,构建连贯事件序列而非孤立帧;(3) 采用GPT-4o驱动的代理引导查询分解,自动解析模糊查询,拆分为视觉/OCR/ASR子查询,并实现自适应得分融合,消除人工模态选择。定性分析表明,该系统能有效应对模糊查询,检索出时序连贯序列,并动态调整融合策略,显著提升交互式片段搜索能力。

原文摘要 · Abstract (English)

The exponential growth of video content has created an urgent need for efficient multimodal moment retrieval systems. However, existing approaches face three critical challenges: (1) fixed-weight fusion strategies fail across cross modal noise and ambiguous queries, (2) temporal modeling struggles to capture coherent event sequences while penalizing unrealistic gaps, and (3) systems require manual modality selection, reducing usability. We propose a unified multimodal moment retrieval system with three key innovations. First, a cascaded dual-embedding pipeline combines BEIT-3 and SigLIP for broad retrieval, refined by BLIP-2 based reranking to balance recall and precision. Second, a temporal-aware scoring mechanism applies exponential decay penalties to large temporal gaps via beam search, constructing coherent event sequences rather than isolated frames. Third, Agent-guided query decomposition (GPT-4o) automatically interprets ambiguous queries, decomposes them into modality specific sub-queries (visual/OCR/ASR), and performs adaptive score fusion eliminating manual modality selection. Qualitative analysis demonstrates that our system effectively handles ambiguous queries, retrieves temporally coherent sequences, and dynamically adapts fusion strategies, advancing interactive moment search capabilities.

视频检索多模态时序建模智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。