arXiv:2604.14605cs.CV2026-04中稿 · CVPR被引 1

让不搭的视觉元素自动匹配风格,生成更和谐的设计。

Towards Design Compositing

论文配图:Towards Design Compositing
图 1 · 摘自论文原文
  • 无需训练,保留原图身份的同时统一风格
  • 在两个设计流水线中显著提升视觉和谐度
  • 适合需要快速融合多源素材的设计师

图形设计创作需将来自不同来源的图像、文字、标志等多模态元素和谐组合成美观一致的整体。现有方法多聚焦于布局预测或补全元素生成,但通常保持输入元素不变,隐含假设其风格已协调。现实中输入元素常因来源不同而存在视觉冲突,该假设限制了实际应用。本文提出 GIST,一种无需训练、保留图像身份的风格统一流水线模块,可无缝嵌入任意现有组件到设计或设计优化流程中。通过与 LaDeCo 和 Design-o-meter 的集成验证,GIST 在视觉和谐性与美学质量上均有显著提升,经 LLaVA-OV 与 GPT-4V 的逐项评分和配对偏好测试确认。

原文摘要 · Abstract (English)

Graphic design creation involves harmoniously assembling multimodal components such as images, text, logos, and other visual assets collected from diverse sources, into a visually-appealing and cohesive design. Recent methods have largely focused on layout prediction or complementary element generation, while retaining input elements exactly, implicitly assuming that provided components are already stylistically harmonious. In practice, inputs often come from disparate sources and exhibit visual mismatch, making this assumption limiting. We argue that identity-preserving stylization and compositing of input elements is a critical missing ingredient for truly harmonized components-to-design pipelines. To this end, we propose GIST, a training-free, identity-preserving image compositor that sits between layout prediction and typography generation, and can be plugged into any existing components-to-design or design-refining pipeline without modification. We demonstrate this by integrating GIST with two substantially different existing methods, LaDeCo and Design-o-meter. GIST shows significant improvements in visual harmony and aesthetic quality across both pipelines, as validated by LLaVA-OV and GPT-4V on aspect-wise ratings and pairwise preference over naive pasting. Project Page: abhinav-mahajan10.github.io/GIST/.

图像合成风格迁移设计自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。