对比扩散与视觉自回归模型的图文组合对齐能力,发现大模型表现更优。
Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
- 首次系统比较扩散与视觉自回归模型在组合对齐上的表现。
- Infinity-8B 在多类任务中表现最佳,2B版也超越部分大模型。
- 适合关注图文生成质量与模型效率的研究者参考。
实现文本描述与生成图像在物体、属性和空间关系上的组合对齐,仍是现代文本到图像(T2I)模型的核心挑战。尽管基于扩散的架构已被广泛研究,但新兴的视觉自回归(VAR)模型的组合行为仍缺乏系统评估。本文在 T2I-CompBench++ 和 GenEval 全套基准上,对六种不同 T2I 系统——SDXL、PixArt-α、Flux-Dev、Flux-Schnell、Infinity-2B 和 Infinity-8B——进行了评测,涵盖颜色与属性绑定、空间关系、计数能力及复杂多对象提示下的对齐表现。在两个基准上,Infinity-8B 均取得最强的整体组合对齐性能;而 Infinity-2B 在多个类别中表现与甚至超过更大的扩散模型,展现出良好的效率-性能权衡。相比之下,SDXL 与 PixArt-α 在属性敏感和空间任务中持续存在短板。这些结果首次系统比较了 VAR 与扩散方法在组合对齐上的差异,并为未来 T2I 模型的发展建立了统一基准。
原文摘要 · Abstract (English)
Achieving compositional alignment between textual descriptions and generated images - covering objects, attributes, and spatial relationships - remains a core challenge for modern text-to-image (T2I) models. Although diffusion-based architectures have been widely studied, the compositional behavior of emerging Visual Autoregressive (VAR) models is still largely unexamined. We benchmark six diverse T2I systems - SDXL, PixArt-$α$, Flux-Dev, Flux-Schnell, Infinity-2B, and Infinity-8B - across the full T2I-CompBench++ and GenEval suites, evaluating alignment in color and attribute binding, spatial relations, numeracy, and complex multi-object prompts. Across both benchmarks, Infinity-8B achieves the strongest overall compositional alignment, while Infinity-2B also matches or exceeds larger diffusion models in several categories, highlighting favorable efficiency-performance trade-offs. In contrast, SDXL and PixArt-$α$ show persistent weaknesses in attribute-sensitive and spatial tasks. These results provide the first systematic comparison of VAR and diffusion approaches to compositional alignment and establish unified baselines for the future development of the T2I model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。