arXiv:2504.04572cs.CV2025-04被引 1

提出多模态长视频检索框架与新评估方法,提升长视频精准定位能力。

Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric

  • 融合视觉、音频与字幕的多模态匹配机制,支持未见词汇和场景检索
  • 在YouCook2数据集上实现良好检索性能,长视频处理效果优于基线
  • 专为长视频设计新评估指标,推动该领域研究发展

精确视频检索需利用多模态关联以应对未见词汇与场景,对长视频而言更具挑战性,因模型须在未预训练特定数据集的情况下仍有效运行。本文提出统一框架,结合视觉匹配流与音频匹配流,并引入基于字幕的视频分段方法;音频流采用两阶段互补检索机制,显著提升长视频检索表现。针对长视频检索的复杂性及其评估难题,提出专用于长视频检索的新评估方法,以支持后续研究。在YouCook2基准上进行实验,结果表明该框架具备出色的检索性能。

原文摘要 · Abstract (English)

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.

视频检索多模态长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。