arXiv:2510.27571cs.CVcs.AI2025-10被引 8

构建通用视频检索新框架,突破传统单一任务局限。

Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum

  • 设计16个数据集组成的诊断性基准UVRB,定位能力短板。
  • 合成155万高质量图文对,覆盖多样化语义空间。
  • 用多模态金字塔课程训练模型,实现零样本泛化领先。

当前视频检索范式结构失配,狭窄的基准测试导致数据和训练任务受限,抑制了通用能力的发展。为此,我们提出一个评估、数据与建模协同设计的框架。首先,建立通用视频检索基准(UVRB),包含16个数据集,不仅衡量性能,还能诊断跨任务、跨领域的关键能力缺口。其次,基于UVRB诊断结果,开发可扩展的合成流程,生成155万条高质量图文对,填补实现通用性所需的语义空间。最后,提出多模态金字塔课程,通过显式利用多样化数据中的潜在关联,训练通用视频嵌入模型(GVE)。大量实验表明,GVE在UVRB上实现零样本泛化性能最优。分析显示,现有主流基准无法有效预测通用能力,且部分相关检索是主导但被忽视的场景。整体框架为突破有限范围、迈向真正通用视频检索提供了可行路径。

原文摘要 · Abstract (English)

The prevailing video retrieval paradigm is structurally misaligned, as narrow benchmarks incentivize correspondingly limited data and single-task training. Therefore, universal capability is suppressed due to the absence of a diagnostic evaluation that defines and demands multi-dimensional generalization. To break this cycle, we introduce a framework built on the co-design of evaluation, data, and modeling. First, we establish the Universal Video Retrieval Benchmark (UVRB), a suite of 16 datasets designed not only to measure performance but also to diagnose critical capability gaps across tasks and domains. Second, guided by UVRB's diagnostics, we introduce a scalable synthesis workflow that generates 1.55 million high-quality pairs to populate the semantic space required for universality. Finally, we devise the Modality Pyramid, a curriculum that trains our General Video Embedder (GVE) by explicitly leveraging the latent interconnections within our diverse data. Extensive experiments show GVE achieves state-of-the-art zero-shot generalization on UVRB. In particular, our analysis reveals that popular benchmarks are poor predictors of general ability and that partially relevant retrieval is a dominant but overlooked scenario. Overall, our co-designed framework provides a practical path to escape the limited scope and advance toward truly universal video retrieval.

视频检索通用模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。