通过分层分解图像细节,实现可控且可解释的图像生成
CART: Compositional Auto-Regressive Transformer for Image Generation
- 将图像分步分解为语义清晰的视觉层,逐步添加细节
- 在三种分解策略下均优于传统方法,提升可控性与分辨率扩展能力
- 适合需要精细控制和物理意义建模的图像生成任务
我们提出一种新型自回归图像生成方法CART,将图像建模为可解释视觉层的层次化组合。尽管自回归模型在自然语言处理中取得突破,但图像固有的空间依赖性使其在视觉任务中难以复制类似成功。CART通过语义明确的分解方式,迭代添加图像细节。我们在三种不同分解策略上验证其通用性:(i) 基础-细节分解(基于Mumford-Shah平滑性),(ii) 内在分解(反照率/光照),(iii) 高光分解(漫反射/镜面反射)。该逐层细化策略超越传统逐词或逐尺度生成方式,在可控性、语义可解释性和分辨率扩展性方面表现更优。实验表明,CART生成图像质量高,并支持结构化编辑,为基于物理或感知动机的图像因子分解提供了新方向。
原文摘要 · Abstract (English)
We propose a novel Auto-Regressive (AR) image generation approach that models images as hierarchical compositions of interpretable visual layers. While AR models have achieved transformative success in language modeling, replicating this success in vision tasks remains challenging due to inherent spatial dependencies in images. Addressing the unique challenges of vision tasks, our method (CART) adds image details iteratively via semantically meaningful decompositions. We demonstrate the flexibility and generality of CART by applying it across three distinct decomposition strategies: (i) Base-Detail Decomposition (Mumford-Shah smoothness), (ii) Intrinsic Decomposition (albedo/shading), and (iii) Specularity Decomposition (diffuse/specular). This next-detail strategy outperforms traditional next-token and next-scale approaches, improving controllability, semantic interpretability, and resolution scalability. Experiments show CART generates visually compelling results while enabling structured image manipulation, opening new directions for controllable generative modeling via physically or perceptually motivated image factorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。