评测视频生成模型对物体状态变化的理解能力,发现现有模型表现不佳。
OSCBench: Benchmarking Object State Change in Text-to-Video Generation

- 基于烹饪指令数据构建三类场景测试物体状态变化
- 六款模型在新组合场景中状态变化准确率不足40%
- 适合关注视频生成理解力与泛化能力的研究者
文本到视频(T2V)生成模型在视觉质量与时间连贯性上进展迅速,但现有评估基准主要关注感知质量、文本-视频对齐或物理合理性,忽略了动作理解中的关键环节:文本提示中明确指定的物体状态变化(OSC)。OSC指动作引发的物体状态转变,如削土豆或切柠檬。本文提出OSCBench,一个专门评估T2V模型在物体状态变化方面表现的基准。OSCBench源自教学类烹饪数据,系统组织出常规、新颖和组合三种场景,以检验模型在分布内性能与泛化能力。我们采用人工评估与多模态大语言模型(MLLM)自动评估两种方式,测试了六款代表性开源及专有T2V模型。结果显示,尽管在语义与场景对齐上表现良好,当前模型在新场景与组合场景中仍普遍存在状态变化不准确且时间不一致的问题,尤其在新异设置下准确率低于40%。该结果表明,物体状态变化是制约文本到视频生成的关键瓶颈,并确立了OSCBench作为诊断性基准,推动具备状态感知能力的视频生成模型发展。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text-video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object's state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action-object interactions into regular, novel, and compositional scenarios to probe both in-distribution performance and generalization. We evaluate six representative open-source and proprietary T2V models using both human user study and multimodal large language model (MLLM)-based automatic evaluation. Our results show that, despite strong performance on semantic and scene alignment, current T2V models consistently struggle with accurate and temporally consistent object state changes, especially in novel and compositional settings. These findings position OSC as a key bottleneck in text-to-video generation and establish OSCBench as a diagnostic benchmark for advancing state-aware video generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。