arXiv:2601.14951cs.CVcs.AI2026-01Conference of the …

首个评估文本生成图像模型时间认知能力的数据集

TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models

  • 构建7.9千条带时间描述的提示数据集,覆盖5类时间知识
  • 人类评估显示各模型时间准确率均低于75%
  • 现有自动评估方法无法可靠判断时间信息理解能力

时间会改变世界中物体、地点和动物的视觉外观。因此,在生成语境相关图像时,对时间的知识与推理至关重要(例如生成春季或冬季的风景)。尽管自然语言处理领域已有大量关于时间知识的研究,但文本到图像(T2I)模型如何呈现和处理时间现象仍鲜有研究。本文提出TempViz,首个全面评估图像生成中时间知识的数据集,包含7.9千条提示和600多张参考图像。利用TempViz,我们评估了五种T2I模型在五个时间知识类别中的表现。人工评估显示,时间理解能力普遍较弱,无一模型在所有类别中准确率超过75%。为进一步开展大规模研究,我们还比较了多种现有自动化评估方法与人工判断的一致性。然而,这些方法均未能可靠评估时间线索——凸显未来在T2I模型时间知识研究上的迫切需求。

原文摘要 · Abstract (English)

Time alters the visual appearance of entities in our world, like objects, places, and animals. Thus, for accurately generating contextually-relevant images, knowledge and reasoning about time can be crucial (e.g., for generating a landscape in spring vs. in winter). Yet, although substantial work exists on understanding and improving temporal knowledge in natural language processing, research on how temporal phenomena appear and are handled in text-to-image (T2I) models remains scarce. We address this gap with TempViz, the first data set to holistically evaluate temporal knowledge in image generation, consisting of 7.9k prompts and more than 600 reference images. Using TempViz, we study the capabilities of five T2I models across five temporal knowledge categories. Human evaluation shows that temporal competence is generally weak, with no model exceeding 75% accuracy across categories. Towards larger-scale studies, we also examine automated evaluation methods, comparing several established approaches against human judgments. However, none of these approaches provides a reliable assessment of temporal cues - further indicating the pressing need for future research on temporal knowledge in T2I.

文本生成图像时间理解评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。