用合成数据训练可让图像分层分解更高效,且避免真实数据的不平衡问题。
Does Synthetic Layered Design Data Benefit Layered Design Decomposition?

- 基于视觉语言模型生成合成分层数据,自动构建训练集
- 50K样本后性能趋于饱和,纯合成数据效果优于真实数据集
- 能均衡控制图层数分布,适合规模化设计编辑系统
近期图像生成技术虽能产出高质量图像,但其结果为扁平化结构,前景、背景与文字混杂于固定画布中,导致生成后编辑困难,限制了实际应用。现有方法依赖稀缺专有分层资源或基于有限结构先验构建部分合成数据,均面临可扩展性瓶颈。本文研究纯合成分层数据是否能提升图形设计分解效果。假设在设计中,分层解耦无需像自然图像那样精确建模层间依赖,因设计元素常为模块化、语义独立组件。我们以当前领先框架CLD为基础,构建合成数据集SynLayers,利用视觉语言模型生成文本监督,并通过VLM预测边界框自动化推理输入。实验发现:(1)仅使用合成数据训练即可超越难以扩展的PrismLayersPro等现有数据集,证明其作为可扩展替代方案的有效性;(2)性能随数据量增加持续提升,约50K样本后趋于饱和;(3)合成数据能实现图层数分布的均衡控制,避免真实数据中常见的层数失衡问题。本研究倡导以合成数据为基石推动分层设计编辑系统的实用化。
原文摘要 · Abstract (English)
Recent advances in image generation have made it easy to produce high-quality images. However, these outputs are inherently flattened, entangling foreground elements, background, and text within a fixed canvas. As a result, flexible post-generation editing remains challenging, revealing a clear last-mile gap toward practical usability. Existing approaches either rely on scarce proprietary layered assets or construct partially synthetic data from limited structural priors. However, both strategies face fundamental challenges in scalability. In this work, we investigate whether pure synthetic layered data can improve graphic design decomposition. We make the assumption that, in graphic design, effective decomposition does not require modeling inter-layer dependencies as precisely as in natural-image composition, since design elements are often intentionally arranged as modular and semantically separable components. Concretely, we conduct a data-centric study based on CLD baseline, which is a state-of-the-art layer decomposition framework. Based on the baseline, we construct our own synthetic dataset, SynLayers, generate textual supervision using vision language models, and automate inference inputs with VLM-predicted bounding boxes. Our study reveals three key findings: (1) even training with purely synthetic data can outperform non-scalable alternatives such as the widely used PrismLayersPro dataset, demonstrating its viability as a scalable and effective substitute; (2) performance consistently improves with increased training data scale, while gains begin to saturate at around 50K samples; and (3) synthetic data enables balanced control over layer-count distributions, avoiding the layer-count imbalance commonly observed in real-world datasets. We hope this data-centric study encourages broader adoption of synthetic data as a practical foundation for layered design editing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。