让AI更懂设计约束,编辑效率提升超四成
StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation

- 将设计编辑转为参数调整,用可微分渲染器传递约束误差
- 约束满足率89%(对比GPT-4V的52%),字体识别准确率76%
- 适合需要高精度排版与视觉一致性的专业设计场景
图形设计编辑需在严格设计约束下精确操控字体、布局与视觉层次。尽管大语言模型推动了视觉语言模型的发展,但现有模型基于像素操作,在结构化设计编辑中仅能达到52%的约束满足率,难以支撑专业工作流。本文提出StructuredEdit,将设计编辑重构为参数调控而非像素生成。核心贡献是可微分参数传播(DPP)训练方法,通过轻量级可微分光栅化器将像素级约束违规反向传播至视觉语言模型微调过程。构建了包含12.5万组验证编辑三元组的混合候选-筛选流水线。系统实现89%的约束满足率(对比GPT-4V的52%)、0.82的匹配元素交并比,以及对前100种常见字体76%的首选准确率。用户研究(N=35)显示,编辑时间减少33%,修正迭代次数下降44%。
原文摘要 · Abstract (English)
Graphic design editing requires precise manipulation of typography, layout, and visual hierarchy under strict design constraints. Following the introduction of large language models, organizations have increasingly promoted vision-language models to enhance productivity. However, current models operate on pixels and achieve only 52% constraint satisfaction on structured design edits, thereby limiting their reliability for professional workflows. We present StructuredEdit, a pipeline that reframes design editing as parameter manipulation rather than pixel generation. Our core technical contribution is Differentiable Parameter Propagation (DPP), a training method that embeds hard design constraints into vision-language model fine-tuning by backpropagating pixel-level constraint violations through a lightweight differentiable rasterizer. A hybrid candidate-and-filter pipeline produces 125k validated edit triplets. The resulting system reaches 89% constraint satisfaction versus 52% for GPT-4V, 0.82 matched-element Intersection over Union, and 76% top-1 font accuracy over the 100 most-frequent design typefaces. In a user study (N=35), editing time drops 33% and correction iterations drop 44% relative to a GPT-4V baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。