从生成图像转向构建有因果逻辑的智能视觉世界
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

- 提出五级演化框架,从原子生成到世界建模
- 现有模型在空间推理和因果理解上仍严重不足
- 适合关注下一代AI视觉系统的研究者与开发者
近期视觉生成模型在逼真度、文字渲染、指令遵循和交互编辑方面取得显著进展,但在空间推理、状态持久性、长时一致性及因果理解方面仍存在明显短板。我们主张,领域应从外观合成转向智能视觉生成:生成结果需基于结构、动态、领域知识和因果关系。为此,提出五级分类体系:原子生成、条件生成、上下文生成、代理生成和世界建模生成,呈现从被动渲染到主动、自主、世界感知生成的演进路径。分析了关键驱动技术,包括流匹配、统一理解和生成模型、改进视觉表征、后训练、奖励建模、数据整理、合成数据蒸馏及采样加速。指出当前评估常因过度强调感知质量而高估进展,忽略结构性、时序性和因果性缺陷。通过基准综述、真实场景压力测试与专家约束案例研究,本路线图提供以能力为中心的视角,用于理解、评估和推进下一代智能视觉生成系统。
原文摘要 · Abstract (English)
Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning, persistent state, long-horizon consistency, and causal understanding. We argue that the field should move beyond appearance synthesis toward intelligent visual generation: plausible visuals grounded in structure, dynamics, domain knowledge, and causal relations. To frame this shift, we introduce a five-level taxonomy: Atomic Generation, Conditional Generation, In-Context Generation, Agentic Generation, and World-Modeling Generation, progressing from passive renderers to interactive, agentic, world-aware generators. We analyze key technical drivers, including flow matching, unified understanding-and-generation models, improved visual representations, post-training, reward modeling, data curation, synthetic data distillation, and sampling acceleration. We further show that current evaluations often overestimate progress by emphasizing perceptual quality while missing structural, temporal, and causal failures. By combining benchmark review, in-the-wild stress tests, and expert-constrained case studies, this roadmap offers a capability-centered lens for understanding, evaluating, and advancing the next generation of intelligent visual generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。