arXiv:2606.10620cs.CVcs.AI2026-06

用四帧连贯图像测试模型对时间变化的想象能力。

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

论文配图:Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency
图 1 · 摘自论文原文
  • 设计四帧序列生成任务,检验模型跨时序一致性
  • 现有模型在保持物体身份与因果顺序上普遍表现不佳
  • 适合评估视频生成、编辑等需要时序理解的任务

当前图像生成模型虽能产出高质量静态图像,但对视觉世界随时间演变的理解仍不清晰。实际应用如分镜设计、分步插图、参考图编辑和视频预览,要求模型在多个视觉状态间保持身份、物体、空间关系和因果顺序的一致性。现有评估多聚焦单图正确性或视频质量,未能检验模型是否能连贯地想象时间序列过程。我们提出 ImageTime,一个以时空一致性为行为探针的诊断基准。给定动作指令,可选参考图指定初始状态,模型需生成包含四个有序关键帧的图像:初始状态、动作起始、过渡状态和最终状态。该四帧协议比单图生成更具时间挑战性,同时避免密集视频动态的干扰。ImageTime 按能力递进组织任务,将每个场景分解为阶段状态谓词、跨帧时间约束和禁止的因果违规。GPT-5.5 在结构化 VLM-as-judge 协议下评分,生成可解释的能力分、子分数和失败标签。多家族基准测试揭示了当前图像生成系统在维持时序连贯视觉状态上的成功、失败与偏差。

原文摘要 · Abstract (English)

Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood. Practical workflows such as storyboarding, step-by-step illustration, reference-guided editing, and video previsualization require models to preserve identities, objects, spatial relations, and causal order across multiple visual states. Existing evaluations largely measure single-image correctness, compositional alignment, or video quality, leaving open whether an image model can coherently imagine a temporally ordered process. We introduce ImageTime, a diagnostic benchmark that uses spatiotemporal consistency as a behavioral probe of visual world modeling in image generation. Given an action instruction, and optionally a reference image specifying the initial state, a model must generate one image containing four ordered key states: initial state, action onset, transition state, and final state. This four-keyframe protocol is more temporally demanding than single-image generation while avoiding the confounds of dense video dynamics. ImageTime organizes tasks with a progressive capability hierarchy and decomposes each scenario into stage-wise state predicates, cross-frame temporal constraints, and forbidden causal violations. GPT-5.5 scores all generated images under a structured VLM-as-judge protocol, producing interpretable capability scores, diagnostic subscores, and failure labels. Through multi-family benchmarking, ImageTime reveals where current image generation systems succeed, fail, and drift when asked to maintain coherent visual world states over time.

图像生成时序建模基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。