arXiv:2501.10692cs.CV2025-01中稿 · ICME 2024被引 10

融合多模态视觉信号与分层文本理解,提升视频片段定位与亮点检测精度。

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

  • 引入RGB、光流、深度图三模态融合机制,动态整合互补视觉信息。
  • 在QVHighlights数据集上MR-mAP@Avg提升3.41,HD-HIT@1提升3.46。
  • 适合关注视频理解、多模态学习与交互式检索的研究者。

给定一段视频和自然语言查询,视频片段定位与亮点检测(MR&HD)旨在定位所有相关时间段并预测其重要性得分。现有方法大多仅使用RGB图像作为输入,忽略了光流、深度图等内在的多模态视觉信号。本文提出多模态融合与查询优化网络(MRNet),通过设计多模态融合模块,动态结合RGB、光流和深度图信息;同时引入查询优化模块,从词、短语、句子三个粒度层次融合文本内容,模拟人类对句子的理解过程。在QVHighlights和Charades数据集上的大量实验表明,MRNet优于当前最优方法,在QVHighlights上达到MR-mAP@Avg +3.41、HD-HIT@1 +3.46的显著提升。

原文摘要 · Abstract (English)

Given a video and a linguistic query, video moment retrieval and highlight detection (MR&HD) aim to locate all the relevant spans while simultaneously predicting saliency scores. Most existing methods utilize RGB images as input, overlooking the inherent multi-modal visual signals like optical flow and depth. In this paper, we propose a Multi-modal Fusion and Query Refinement Network (MRNet) to learn complementary information from multi-modal cues. Specifically, we design a multi-modal fusion module to dynamically combine RGB, optical flow, and depth map. Furthermore, to simulate human understanding of sentences, we introduce a query refinement module that merges text at different granularities, containing word-, phrase-, and sentence-wise levels. Comprehensive experiments on QVHighlights and Charades datasets indicate that MRNet outperforms current state-of-the-art methods, achieving notable improvements in MR-mAP@Avg (+3.41) and HD-HIT@1 (+3.46) on QVHighlights.

视频理解多模态检索亮点检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。