通过持续追踪语义承诺,提升复杂图像生成的准确性。
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation

- 构建结构化规范动态维护语义要求,实现全程追踪。
- 在Gen-Arena上达0.60的实体优先通过率,显著优于基线。
- 适合需要精确控制复杂视觉意图的生成任务研究者。
尽管文本到图像模型在视觉保真度上取得显著进展,但忠实实现复杂视觉意图仍具挑战性,因多个要求需贯穿定位、生成与验证全过程。我们称这些要求为语义承诺,并将承诺生命周期断裂现象定义为概念裂隙——承诺可能局部解决或检查,但在生成过程中无法保持为同一操作单元。为此,我们提出SCOPE:一种基于规范引导的技能调度框架,通过动态演化结构化规范持续维护语义承诺,并在未解决或被违反时条件性调用检索、推理与修复技能。为评估承诺级意图实现,我们引入Gen-Arena,一个包含实体与约束级规范的人工标注基准,以及以实体优先为标准的实体门控意图通过率(EGIP)。SCOPE在Gen-Arena上达到0.60的EGIP,同时在WISE-V(0.907)和MindBench(0.61)上也表现优异,证明持续追踪语义承诺对复杂图像生成的有效性。
原文摘要 · Abstract (English)
While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many requirements must be tracked across grounding, generation, and verification. We refer to these requirements as semantic commitments and formalize their lifecycle discontinuity as the Conceptual Rift, where commitments may be locally resolved or checked but fail to remain identifiable as the same operational units throughout the generation lifecycle. To address this, we propose SCOPE, a specification-guided skill orchestration framework that maintains semantic commitments in an evolving structured specification and conditionally invokes retrieval, reasoning, and repair skills around unresolved or violated commitments. To evaluate commitment-level intent realization, we introduce Gen-Arena, a human-annotated benchmark with entity- and constraint-level specifications, together with Entity-Gated Intent Pass Rate (EGIP), a strict entity-first pass criterion. SCOPE substantially outperforms all evaluated baselines on Gen-Arena, achieving 0.60 EGIP, and further achieves strong results on WISE-V (0.907) and MindBench (0.61), demonstrating the effectiveness of persistent commitment tracking for complex image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。