让大模型学会可解释的空间推理,自动设计视觉布局。
From Pixels to Policies: Reinforcing Spatial Reasoning in Language Models for Content-Aware Layout Design
- 将布局设计转为文本化空间环境下的策略学习,显式建模元素位置与关系。
- 生成可追踪的推理过程和结构化布局,提升设计透明度与可控性。
- 在少样本下超越大型闭源模型,适合需要可解释性的智能设计场景。
我们提出LaySPA,一种强化学习框架,使大语言模型具备显式且可解释的空间推理能力,用于内容感知的图形布局设计。该框架解决两大挑战:大模型空间推理能力有限,以及设计决策过程不透明。不同于像素级操作,我们将在结构化文本空间环境中进行布局设计策略学习,显式编码画布几何、元素属性及元素间关系。LaySPA输出双层结果:可解释的推理轨迹与结构化布局规范,实现透明且可控的设计决策。通过多目标空间评判,将布局质量分解为几何有效性、关系一致性与美学一致性,并采用相对组优化训练方法,稳定开放设计空间中的学习过程。实验表明,LaySPA在结构有效性和视觉质量上均有提升,优于更大的专有大模型,在性能上接近专用前沿布局生成器,同时所需标注样本更少、延迟更低。
原文摘要 · Abstract (English)
We introduce LaySPA, a reinforcement learning framework that equips large language models (LLMs) with explicit and interpretable spatial reasoning for content-aware graphic layout design. LaySPA addresses two key challenges: LLMs' limited spatial reasoning and the lack of opacity in design decision making. Instead of operating at the pixel level, we reformulate layout design as a policy learning problem over a structured textual spatial environment that explicitly encodes canvas geometry, element attributes, and inter-element relationships. LaySPA produces dual-level outputs comprising interpretable reasoning traces and structured layout specifications, enabling transparent and controllable design decision making. Layout design policy is optimized via a multi-objective spatial critique that decomposes layout quality into geometric validity, relational coherence, and aesthetic consistency, and is trained using relative group optimization to stabilize learning in open-ended design spaces. Experiments demonstrate that LaySPA improves structural validity and visual quality, outperforming larger proprietary LLMs and achieving performance comparable to specialized SOTA layout generators while requiring fewer annotated samples and reduced latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。