arXiv:2609.02864cs.CV2026-09

评测视觉生成模型的推理能力,发现现有模型常生成局部合理但全局错误的结果。

Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation

论文配图:Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation
图 1 · 摘自论文原文
  • 构建四类认知挑战任务,系统评估生成模型的推理-生成能力。
  • 2000个样本测试显示主流模型存在显著推理与生成差距。
  • 适合研究生成式AI逻辑一致性与下一代世界模型的开发者使用。

统一生成模型(UGMs)和世界模拟器在视觉感知与合成方面取得了突破性进展,但主要依赖表面事件对齐,高阶视觉推理能力仍未被充分探索。真正的视觉生成智能需要「推理到生成」能力,即从视觉输入中推断潜在规则,并以逻辑严谨、精确约束的方式呈现结果。本文提出RIG-BENCH,一个全新的综合性基准,系统评估四类认知挑战领域下的推理驱动图像生成(RIG):概念类、变换类、模式与结构类、场景类。该基准包含2000个精心筛选的样本,可作为对RIG能力的严格压力测试。我们对前沿UGMs及图像/视频生成模型的广泛评估揭示了显著的推理-生成差距:模型常生成局部合理但全局不合逻辑的输出。RIG-BENCH为下一代逻辑严谨的UGMs与世界模拟器的发展提供了关键诊断框架。

原文摘要 · Abstract (English)

Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intelligence demands "Reasoning-to-Generation", an ability to infer latent rules from visual inputs and manifest solutions through precise, logically constrained visual outcomes. We introduce RIG-BENCH, a novel comprehensive benchmark that systematically evaluates Reasoning-driven Image Generation (RIG) across four cognitively demanding domains: Concept-based, Transformation-based, Pattern & Structure, and Scenario-based. Featuring 2000 curated samples, RIG-BENCH serves as a rigorous stress test for RIG. Our extensive evaluations of state-of-the-art UGMs and image/video generation models reveal a significant reasoning-generation gap, wherein models frequently produce locally plausible but globally illogical outputs. RIG-BENCH provides a vital diagnostic framework to guide the development of next-generation, logically grounded UGMs and world simulators.

视觉生成推理能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。