构建155万张分层设计图数据集,支持结构化图文生成与编辑研究。
LICA: Layered Image Composition Annotations for Graphic Design Research
- 将设计拆解为带元数据的分层组件,支持精确结构操作。
- 包含27,261个动画布局及关键帧,覆盖97万种模板。
- 适合研究结构化设计生成、可控编辑与时间感知建模者。
我们提出LICA(Layered Image Composition Annotations),一个包含1,550,244个多层图形设计作品的大规模数据集,旨在推动对图形布局的结构化理解与生成。除渲染后的PNG图像外,LICA以层级结构表示每个设计,包含文本、图像、矢量和组元素等类型组件,并为每项元素配备丰富的元数据,如空间几何、字体属性、不透明度和可见性。数据集涵盖20类设计,包含971,850个唯一模板,覆盖真实世界的设计结构。我们进一步引入图形设计视频这一全新挑战,标注了27,261个动态布局,含每组件的关键帧与运动参数。除规模外,LICA确立了图形设计研究的新范式,支持层感知修复、结构化布局生成、受控设计编辑和时间感知生成建模等任务。通过将设计视为组件与关系的系统,该数据集支持直接作用于设计结构而非仅像素的模型研究。
原文摘要 · Abstract (English)
We introduce LICA (Layered Image Composition Annotations), a large scale dataset of 1,550,244 multi-layer graphic design compositions designed to advance structured understanding and generation of graphic layouts. In addition to rendered PNG images, LICA represents each design as a hierarchical composition of typed components including text, image, vector, and group elements, each paired with rich per-element metadata such as spatial geometry, typographic attributes, opacity, and visibility. The dataset spans 20 design categories and 971,850 unique templates, providing broad coverage of real-world design structures. We further introduce graphic design video as a new and largely unexplored challenge for current vision-language models through 27,261 animated layouts annotated with per-component keyframes and motion parameters. Beyond scale, LICA establishes a new paradigm of research tasks for graphic design, enabling structured investigations into problems such as layer-aware inpainting, structured layout generation, controlled design editing, and temporally-aware generative modeling. By representing design as a system of compositional layers and relationships, the dataset supports research on models that operate directly on design structure rather than pixels alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。