构建动态因果生成基准,检验模型对世界过程的理解与生成能力
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- 设计四阶段因果事件生成任务,覆盖六大学科领域
- 统一模型在因果连贯性上优于专用图像生成模型
- 提出综合评估指标Envision-Score,聚焦时序一致性与物理合理性
当前多模态模型试图通过统一理解与生成来突破单模态表示的局限,常以文本到图像(T2I)任务校准语义一致性。然而,其在训练与评估中依赖静态单图生成,导致过度拟合静态模式匹配和语义融合,根本阻碍了对随时间演化的动态过程建模。为此,我们提出Envision——一个针对链式文本到多图生成的因果事件演化基准。该基准基于世界知识,以时空因果结构组织,重构现有评估维度,包含1000个四阶段提示,覆盖六个科学与人文领域。为实现从单图到序列帧的评估转变,并检验模型是否真正内化世界知识并遵守因果-时序约束,我们引入Envision-Score,一个整合多维一致性、物理合理性和美学性的综合指标。对15个模型(10个专用T2I模型,5个统一模型)的全面评估显示:专用T2I模型在美学渲染上表现优异,但缺乏内在世界知识;统一多模态模型弥补这一差距,在因果叙事连贯性上持续优于专用模型。然而,即使这些统一架构仍逊于闭源模型,且难以克服时空一致性的核心挑战。这表明,关注因果隔离的单图生成会抑制多帧推理与生成,促进静态模式匹配而非动态世界建模,最终限制世界知识的内化与生成能力。
原文摘要 · Abstract (English)
Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to calibrate semantic consistency. However, their reliance on static, single-image generation in training and evaluation leads to overfitting to static pattern matching and semantic fusion, while fundamentally hindering their ability to model dynamic processes that unfold over time. To address these constraints, we propose Envision-a causal event progression benchmark for chained text-to-multi-image generation. Grounded in world knowledge and structured by spatiotemporal causality, it reorganizes existing evaluation dimensions and includes 1,000 four-stage prompts spanning six scientific and humanities domains. To transition evaluation from single images to sequential frames and assess whether models truly internalize world knowledge while adhering to causal-temporal constraints, we introduce Envision-Score, a holistic metric integrating multi-dimensional consistency, physicality, and aesthetics. Comprehensive evaluation of 15 models (10 specialized T2I models, 5 unified models) uncovers: specialized T2I models demonstrate proficiency in aesthetic rendering yet lack intrinsic world knowledge. Unified multimodal models bridge this gap, consistently outperforming specialized counterparts in causal narrative coherence. However, even these unified architectures remain subordinate to closed-source models and struggle to overcome the core challenge of spatiotemporal consistency. This demonstrates that a focus on causally-isolated single images impedes multi-frame reasoning and generation, promoting static pattern matching over dynamic world modeling-ultimately limiting world knowledge internalization, generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。