arXiv:2603.14426cs.CVcs.IR2026-03

构建可控状态变化的AI生成视频数据集,评测模型对时序终点的感知能力。

GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos

  • 设计包含时序与语义硬负例的对抗性数据集,区分时间与语义混淆
  • 发现主流模型常误判仅终点状态不同的视频,对语义替换不敏感
  • 提供基于三元组的诊断分析,适合研究时序推理与状态感知的学者

现有文本到视频检索基准多基于真实影像,语义可从单帧推断,导致时序推理与终态定位评估不足。本文提出GenState-AI,一个聚焦可控状态转移的AI生成基准,每个查询配对主视频、仅终态不同的时序硬负例,以及内容替换的语义硬负例,实现对时序与语义混淆的细粒度诊断。基于Wan2.2-TI2V-5B生成短片段,其语义依赖于位置、数量及对象关系的精确变化,提供可控制的评估条件。评估两个代表性多模态大模型基线,发现两者均频繁将主视频误判为时序硬负例,偏好时序合理但终态错误的片段,表明对关键终态证据的建模不足,而对语义替换相对鲁棒。进一步引入三元组诊断分析,包括相对顺序统计与转换类别分解,明确区分失败来源。GenState-AI为状态感知、时空敏感的文本到视频检索提供专注测试平台,将在huggingface.co发布。

原文摘要 · Abstract (English)

Existing text-to-video retrieval benchmarks are dominated by real-world footage where much of the semantics can be inferred from a single frame, leaving temporal reasoning and explicit end-state grounding under-evaluated. We introduce GenState-AI, an AI-generated benchmark centered on controlled state transitions, where each query is paired with a main video, a temporal hard negative that differs only in the decisive end-state, and a semantic hard negative with content substitution, enabling fine-grained diagnosis of temporal vs. semantic confusions beyond appearance matching. Using Wan2.2-TI2V-5B, we generate short clips whose meaning depends on precise changes in position, quantity, and object relations, providing controllable evaluation conditions for state-aware retrieval. We evaluate two representative MLLM-based baselines, and observe consistent and interpretable failure patterns: both frequently confuse the main video with the temporal hard negative and over-prefer temporally plausible but end-state-incorrect clips, indicating insufficient grounding to decisive end-state evidence, while being comparatively less sensitive to semantic substitutions. We further introduce triplet-based diagnostic analyses, including relative-order statistics and breakdowns across transition categories, to make temporal vs. semantic failure sources explicit. GenState-AI provides a focused testbed for state-aware, temporally and semantically sensitive text-to-video retrieval, and will be released on huggingface.co.

视频检索状态感知AI生成时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。