arXiv:2511.10136cs.CVcs.AI2025-11中稿 · AAAI被引 2

现有文生图模型在组合逻辑上严重失效,原因在于训练数据和架构设计缺陷。

Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation

  • 分析否定、计数、空间关系三类逻辑组合的失败机制
  • 组合任务准确率远低于单一任务,表现急剧下降
  • 适合关注生成模型本质缺陷的研究者

当前主流文生图模型的架构存在根本性缺陷:无法有效处理逻辑组合。本文针对否定、计数和空间关系三种核心语义组合进行系统分析,发现当单一逻辑成分组合时,模型性能出现灾难性下降,暴露出严重的干扰问题。我们归因于三点:一是训练数据中几乎不存在显式否定;二是连续注意力架构不适用于离散逻辑;三是评估指标更青睐视觉合理性而非约束满足。通过分析近期基准与方法,表明现有解决方案与简单扩展均无法弥合该差距。结论是,实现真正组合性需在表征与推理层面进行根本性突破,而非对现有架构的渐进优化。

原文摘要 · Abstract (English)

The architectural blueprint of today's leading text-to-image models contains a fundamental flaw: an inability to handle logical composition. This survey investigates this breakdown across three core primitives-negation, counting, and spatial relations. Our analysis reveals a dramatic performance collapse: models that are accurate on single primitives fail precipitously when these are combined, exposing severe interference. We trace this failure to three key factors. First, training data show a near-total absence of explicit negations. Second, continuous attention architectures are fundamentally unsuitable for discrete logic. Third, evaluation metrics reward visual plausibility over constraint satisfaction. By analyzing recent benchmarks and methods, we show that current solutions and simple scaling cannot bridge this gap. Achieving genuine compositionality, we conclude, will require fundamental advances in representation and reasoning rather than incremental adjustments to existing architectures.

文生图组合性逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。