arXiv:2512.04082cs.CV2025-12被引 4

让AI更懂布局逻辑,支持专业级逐层修改海报设计

PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design

  • 分三阶段训练模型,提升几何准确性和审美判断力
  • 生成的海报布局精准且美观,支持多轮精细调整
  • 适合需要反复修改的设计师或广告创意团队

图形设计是现代视觉传播的核心,广泛用于文化与商业活动推广。尽管大型多模态模型(LMMs)已尝试自动化设计流程,但现有方法常产生几何不准确的版面,且缺乏专业工作流所需的迭代式、分层编辑能力。为此,我们提出PosterCopilot框架,通过渐进式三阶段训练策略——扰动监督微调、视觉-现实对齐强化学习、基于美学反馈的强化学习——赋予LMM几何理解与审美推理能力。同时构建完整工作流,结合生成模型实现分层可控、可迭代的编辑,精确调整元素位置并保持整体视觉一致性。大量实验表明,PosterCopilot生成的布局在几何准确性与美学表现上均显著优于现有方法,为专业级设计提供了前所未有的可控性。

原文摘要 · Abstract (English)

Graphic design forms the cornerstone of modern visual communication, serving as a vital medium for promoting cultural and commercial events. Recent advances have explored automating this process using Large Multimodal Models (LMMs), yet existing methods often produce geometrically inaccurate layouts and lack the iterative, layer-specific editing required in professional workflows. To address these limitations, we present PosterCopilot, a framework that advances layout reasoning and controllable editing for professional graphic design. Specifically, we introduce a progressive three-stage training strategy that equips LMMs with geometric understanding and aesthetic reasoning for layout design, consisting of Perturbed Supervised Fine-Tuning, Reinforcement Learning for Visual-Reality Alignment, and Reinforcement Learning from Aesthetic Feedback. Furthermore, we develop a complete workflow that couples the trained LMM-based design model with generative models, enabling layer-controllable, iterative editing for precise element refinement while maintaining global visual consistency. Extensive experiments demonstrate that PosterCopilot achieves geometrically accurate and aesthetically superior layouts, offering unprecedented controllability for professional iterative design.

图像生成设计自动化多模态交互编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。