让图像生成更美:用可迁移的构图控制实现个性化美学创作
Advancing Aesthetic Image Generation via Composition Transfer

- 基于美学理论,提取参考图构图特征,通过条件引导控制生成构图
- 在无参考图时,用大视觉语言模型实现主题驱动的构图检索与生成
- 支持显式构图规划,适合追求艺术美感的创作者和设计应用
构图是视觉美学的核心,虽独立于具体内容,但实际常与语义耦合。现有方法多依赖隐式学习或语义布局控制,缺乏对构图本身的显式建模。为此,我们提出Composer框架,基于美学理论实现语义无关的构图建模。首先,通过提取参考图中的构图感知表征,并结合定制化条件引导模块,控制预训练扩散模型的构图生成。其次,在仅提供文本主题而无构图参考时,利用大视觉语言模型(LVLMs)的上下文学习能力,实现主题驱动的构图检索与生成。此外,通过在训练后的控制模块上进行文本到构图微调,实现无参考模式下的隐式构图规划。我们还构建了一个包含200万张图像-文本对的高质量数据集,以支持模型训练。实验表明,Composer显著提升文本到图像生成任务中的美学质量,支持个性化构图控制与迁移,为创作过程提供精准与灵活性。
原文摘要 · Abstract (English)
Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, in practice, composition is often coupled with semantics. As a result, existing methods often enhance composition either through implicit learning or by semantics-based layout control, rather than explicitly modeling composition itself. To address this gap, we introduce Composer, a framework rooted in aesthetic theory, designed to model composition in a semantic-agnostic manner. First, it supports composition transfer by extracting key composition-aware representations from a reference image and leveraging a tailored conditional guidance module to control composition based on pre-trained diffusion models. Second, when users specify only text themes without a composition reference, Composer supports theme-driven composition retrieval by leveraging the in-context learning capabilities of Large Vision-Language Models (LVLMs), achieving explicit composition planning. To enhance composition in a reference-free mode, we conduct text-to-composition fine-tuning on the trained control module to enable implicit composition planning. Furthermore, we curated a high-quality dataset comprising 2 million image-text pairs using state-of-the-art generative models to support model training. Experimental results demonstrate that Composer significantly enhances aesthetic quality in text-to-image tasks and facilitates personalized composition control and transfer, offering users precision and flexibility in the creative process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。