arXiv:2507.02316cs.CV2025-07被引 7

构建合成视频检索评估基准,验证其在下游任务中的实际价值。

Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos

  • 基于800个真实用户查询生成合成视频,标注四维语义对齐
  • 发现现有质量评估指标与检索性能相关性弱,需新评价体系
  • 提出自动评估工具,助力筛选高价值合成数据,提升检索效果

文本到视频(T2V)合成技术快速发展,但现有评估指标主要关注视觉质量和时序一致性,难以反映合成视频在下游任务如文本到视频检索(TVR)中的实际效用。本文提出SynTVA数据集与评估框架,基于MSRVTT训练集的800个多样化用户查询,使用先进T2V模型生成合成视频,并对每对视频-文本在物体与场景、动作、属性及提示保真度四个关键语义维度进行标注。评估框架分析通用视频质量评估(VQA)指标与对齐得分的相关性,并检验其对下游TVR性能的预测能力。为进一步探索规模化路径,我们开发了自动评估器,可从现有指标推断对齐质量。实验表明,SynTVA可作为高质量数据增强资源,通过筛选高实用性合成样本显著提升TVR性能。项目主页与数据集见https://jasoncodemaker.github.io/SynTVA/。

原文摘要 · Abstract (English)

Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synthetic videos perform in downstream tasks such as text-to-video retrieval (TVR). In this work, we introduce SynTVA, a new dataset and benchmark designed to evaluate the utility of synthetic videos for building retrieval models. Based on 800 diverse user queries derived from MSRVTT training split, we generate synthetic videos using state-of-the-art T2V models and annotate each video-text pair along four key semantic alignment dimensions: Object \& Scene, Action, Attribute, and Prompt Fidelity. Our evaluation framework correlates general video quality assessment (VQA) metrics with these alignment scores, and examines their predictive power for downstream TVR performance. To explore pathways of scaling up, we further develop an Auto-Evaluator to estimate alignment quality from existing metrics. Beyond benchmarking, our results show that SynTVA is a valuable asset for dataset augmentation, enabling the selection of high-utility synthetic samples that measurably improve TVR outcomes. Project page and dataset can be found at https://jasoncodemaker.github.io/SynTVA/.

文本生成视频视频检索数据评估合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。