用推理模型提升视频检索准确率,效果比传统方法强31%。
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
- 基于视觉内容推理判断查询与视频相关性
- 在MultiVENT 2.0上nDCG@10平均提升31%
- 适合需要高精度视频检索的应用场景
重排序是现代检索系统的关键组件,通常将高效的一阶段检索器与更强大的模型结合以优化结果。尽管大语言模型在文本重排序中取得显著进展,但面向视频检索的推理式重排序仍研究不足。为此,我们提出RANKVIDEO,一种基于推理的视频检索重排序器,通过显式分析查询-视频对的视频内容来评估相关性。RANKVIDEO采用两阶段课程训练:先进行感知基础的监督微调,再通过点对点、成对及教师置信度蒸馏目标进行重排序训练,并辅以数据合成流水线生成高推理强度的查询-视频对。在大规模MultiVENT 2.0基准上的实验表明,RANKVIDEO在两阶段框架中持续提升性能,平均提升nDCG@10达31%,优于纯文本与视觉-语言重排序方法,且更具效率。
原文摘要 · Abstract (English)
Reranking is a critical component of modern retrieval systems, which typically pair an efficient first-stage retriever with a more expressive model to refine results. While large reasoning models have driven rapid progress in text-centric reranking, reasoning-based reranking for video retrieval remains underexplored. To address this gap, we introduce RANKVIDEO, a reasoning-based reranker for video retrieval that explicitly reasons over query-video pairs using video content to assess relevance. RANKVIDEO is trained using a two-stage curriculum consisting of perception-grounded supervised fine-tuning followed by reranking training that combines pointwise, pairwise, and teacher confidence distillation objectives, and is supported by a data synthesis pipeline for constructing reasoning-intensive query-video pairs. Experiments on the large-scale MultiVENT 2.0 benchmark demonstrate that RANKVIDEO consistently improves retrieval performance within a two-stage framework, yielding an average improvement of 31% on nDCG@10 and outperforming text-only and vision-language reranking alternatives, while more efficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。