分析文本到视频检索性能瓶颈,发现简单描述更易匹配。
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

- 统一框架评测14种模型,解析查询词特征与性能关系。
- 短而清晰的单动作描述召回率更高,复杂场景仍难处理。
- 适合关注检索系统设计与数据质量的研究者参考。
文本到视频检索使用户能用自然语言查找视频内容,随在线视频激增愈发重要。过去六年涌现双编码器、注意力驱动模型及多模态融合方法,但模型行为、数据集影响和查询难度等基本问题仍不明确。本文在统一预处理与评估框架下,对14种先进检索方法在3个常用数据集上进行评估,分析字幕长度、清晰度、语义类别及动作与场景平衡等特征,并将其与模型表现关联。结果显示,简短、清晰、单一动作或颜色属性的字幕召回率更高;而复杂事件、多步活动或细粒度场景描述对所有模型仍具挑战。注意力驱动架构更擅长处理时序依赖或多步查询,双编码器与多模态融合模型则主要在简单或单类别字幕上表现良好。跨数据集泛化能力随更大更多样字幕集提升,但生成式字幕未显著提高检索准确率。整体揭示了关键数据集因素、基准挑战以及查询内容与模型架构的交互关系,为构建更有效的文本到视频检索系统提供指导。
原文摘要 · Abstract (English)
Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with the rapid expansion of online video. Over the past six years, research has produced numerous methods, such as dual encoders, attention-driven models, and multimodal fusion approaches; however, fundamental questions remain about model behavior, dataset influence, and query difficulty. In this work, we evaluate 14 state-of-the-art retrieval methods across 3 widely used datasets under a unified preprocessing and evaluation framework. We analyze caption characteristics, including length, clarity, semantic category, and Action vs. Scene balance, and link these to model performance. Our results show that short, clear, and simple captions, such as those describing single actions or color attributes, achieve higher recall, while complex events, multi-step activities, or fine-grained scene descriptions remain challenging for all existing models. Attention-driven architectures better handle temporally dependent or multi-step queries, whereas dual-encoder and multimodal fusion models perform well primarily on simpler or single-category captions. Cross-dataset generalization improves with larger, more diverse caption sets, but generative captions do not consistently enhance retrieval accuracy. Overall, our findings highlight key dataset factors, benchmark challenges, and the interplay between query content and model architecture, providing guidance for developing more effective text-to-video retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。