arXiv:2410.12407cs.CVcs.CL2024-10中稿 · ACCV 2024被引 3

提出细粒度评估方法,检测模型对文本细微差异的识别能力。

Beyond Coarse-Grained Matching in Video-Text Retrieval

  • 通过单字替换生成难负样本,实现细粒度评测
  • 现有基准无法有效检验模型对微小差异的感知能力
  • 新基线提升模型对细微语义差别的理解

视频-文本检索虽取得显著进展,但模型对描述中细微差异的辨别能力仍需验证。本文提出一种细粒度评估新方法,可通过自动在名词、动词、形容词、副词和介词等位置进行单字替换,生成具有细微差别的难负样本。我们在两个标准基准(MSR-VTT 和 VATEX)及两个精心构建的详述数据集(VLN-UVO 和 VLN-OOPS)上,对四种前沿模型进行了全面实验,得出若干新发现:1)现有评估基准难以有效检测模型对单字级差异的感知能力;2)细粒度评估揭示了模型在区分此类细微变化时面临巨大挑战。为此,我们提出一个可与现有方法轻松结合的新基线,实验表明该方法显著提升了模型对细粒度语义差异的理解能力。

原文摘要 · Abstract (English)

Video-text retrieval has seen significant advancements, yet the ability of models to discern subtle differences in captions still requires verification. In this paper, we introduce a new approach for fine-grained evaluation. Our approach can be applied to existing datasets by automatically generating hard negative test captions with subtle single-word variations across nouns, verbs, adjectives, adverbs, and prepositions. We perform comprehensive experiments using four state-of-the-art models across two standard benchmarks (MSR-VTT and VATEX) and two specially curated datasets enriched with detailed descriptions (VLN-UVO and VLN-OOPS), resulting in a number of novel insights: 1) our analyses show that the current evaluation benchmarks fall short in detecting a model's ability to perceive subtle single-word differences, 2) our fine-grained evaluation highlights the difficulty models face in distinguishing such subtle variations. To enhance fine-grained understanding, we propose a new baseline that can be easily combined with current methods. Experiments on our fine-grained evaluations demonstrate that this approach enhances a model's ability to understand fine-grained differences.

视频文本检索细粒度评估语义差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。