arXiv:2509.16891cs.AI2025-09被引 1

让大模型学会精准布局,生成视觉平衡的图文设计

LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation

  • 用强化学习增强大模型的空间推理能力
  • 生成的布局结构合理且视觉美观,优于通用大模型
  • 适合需要智能排版的设计师和自动化创作工具

尽管大型语言模型在文本领域展现出强大的推理与规划能力,并能有效执行复杂任务指令,但其对空间关系的理解与操作能力仍有限。这一能力对于内容感知的图形布局设计至关重要,目标是将异构元素合理安排在画布上,确保最终设计视觉平衡且结构可行。该任务需要精确协调多个元素的位置、对齐与结构组织。为此,我们提出LaySPA——一种基于强化学习的框架,通过显式空间推理能力增强基于大模型的智能体。LaySPA采用混合奖励信号,同时捕捉几何约束、结构保真度与视觉质量,使智能体能够导航画布、建模元素间关系并优化空间布局。通过群体相对策略优化,智能体生成反映显著区域、遵守空间约束的内容感知布局,并输出可解释的推理过程与结构化布局规范。实验结果表明,LaySPA显著提升结构有效且视觉吸引人的布局生成效果,优于更大规模通用大模型,性能接近顶尖专用布局模型。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have demonstrated impressive reasoning and planning abilities in textual domains and can effectively follow instructions for complex tasks, their ability to understand and manipulate spatial relationships remains limited. Such capabilities are crucial for content-aware graphic layout design, where the goal is to arrange heterogeneous elements onto a canvas so that final design remains visually balanced and structurally feasible. This problem requires precise coordination of placement, alignment, and structural organization of multiple elements within a constrained visual space. To address this limitation, we introduce LaySPA, a reinforcement learning-based framework that augments LLM-based agents with explicit spatial reasoning capabilities for layout design. LaySPA employs hybrid reward signals that jointly capture geometric constraints, structural fidelity, and visual quality, enabling agents to navigate the canvas, model inter-element relationships, and optimize spatial arrangements. Through group-relative policy optimization, the agent generates content-aware layouts that reflect salient regions, respect spatial constraints, and produces an interpretable reasoning trace explaining placement decisions and a structured layout specification. Experimental results show that LaySPA substantially improves the generation of structurally valid and visually appealing layouts, outperforming larger general-purpose LLMs and achieving performance comparable to state-of-the-art specialized layout models.

布局生成大模型空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。