arXiv:2410.01615cs.CV2024-10被引 29

用视觉显著性引导注意力,提升视频片段检索与高光检测效果

Saliency-Guided DETR for Moment Retrieval and Highlight Detection

  • 引入显著性引导交叉注意力机制,优化文本与视频特征对齐
  • 在QVHighlights等三个基准上达到当前最佳性能
  • 适用于零样本和微调场景,适合实际视频理解应用

现有的视频片段检索与高光检测方法难以高效对齐文本与视频特征,导致性能不佳且难以投入生产。为此,我们提出一种新架构,利用面向对齐的先进基础视频模型。结合提出的显著性引导交叉注意力机制与混合式DETR结构,显著提升了片段检索与高光检测的表现。为进一步优化,我们构建了InterVid-MR——一个大规模高质量预训练数据集。基于该数据集,所提架构在QVHighlights、Charades-STA和TACoS三个基准上均取得领先结果。该方法为视频-语言任务中的零样本与微调场景提供了高效可扩展的解决方案。

原文摘要 · Abstract (English)

Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited production usage. To address this, we propose a novel architecture that utilizes recent foundational video models designed for such alignment. Combined with the introduced Saliency-Guided Cross Attention mechanism and a hybrid DETR architecture, our approach significantly enhances performance in both moment retrieval and highlight detection tasks. For even better improvement, we developed InterVid-MR, a large-scale and high-quality dataset for pretraining. Using it, our architecture achieves state-of-the-art results on the QVHighlights, Charades-STA and TACoS benchmarks. The proposed approach provides an efficient and scalable solution for both zero-shot and fine-tuning scenarios in video-language tasks.

视频检索注意力机制多模态DETR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。