统一图像生成与编辑的推理框架,提升复杂场景理解能力。
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
- 用世界知识增强文本推理,生成时补全隐含信息
- 通过自反思编辑修正视觉错误,实现生成-编辑闭环
- 支持文化常识、物理规律等五大领域,适合高阶图像创作
统一多模态模型在需要深度推理的复杂合成任务中表现不佳,通常将文生图与图像编辑视为独立能力。为此,我们提出UniReason,一种通过两种互补推理范式统一这两项任务的框架。在生成阶段融入增强世界知识的文本推理,以推断隐含信息;在编辑阶段利用细粒度编辑能力进行视觉纠错,实现自反思式修正。该方法在共享架构内整合生成与编辑,模拟人类先规划后优化的认知过程。我们构建了一个大规模(约30万样本)以推理为核心的基准数据集,覆盖文化常识、物理等五大知识领域,并引入代理生成的视觉修正语料。大量实验表明,UniReason在WISE、KrisBench和UniREditBench等推理密集型基准上达到先进水平,同时保持卓越的通用合成能力。
原文摘要 · Abstract (English)
Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoning steps. To address this, we propose UniReason, a unified framework that harmonizes these two tasks through two complementary reasoning paradigms. We incorporate world knowledge-enhanced textual reasoning into generation to infer implicit knowledge, and leverage editing capabilities for fine-grained editing-like visual refinement to further correct visual errors via self-reflection. This approach unifies generation and editing within a shared architecture, mirroring the human cognitive process of planning followed by refinement. We support this framework by systematically constructing a large-scale reasoning-centric dataset (~300k samples) covering five major knowledge domains (e.g., cultural commonsense, physics, etc.) for textual reasoning, alongside an agent-generated corpus for visual refinement. Extensive experiments demonstrate that UniReason achieves advanced performance on reasoning-intensive benchmarks such as WISE, KrisBench and UniREditBench, while maintaining superior general synthesis capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。