arXiv:2505.22523cs.CV2025-05被引 14

开源20万张高精度分层透明图像数据集,推动文本生成可编辑分层图像发展

PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

论文配图:PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
图 1 · 摘自论文原文
  • 用现成扩散模型合成高质量分层透明图像,无需训练
  • 推出ART+模型,视觉效果媲美FLUX.1-[dev],60%用户更偏好
  • 数据集含精确透明度掩码,适合创意设计与可编辑图像应用

从文本提示生成高质量、多层透明图像能带来前所未有的创作控制力,使每层可如编辑文本般轻松修改。然而,由于缺乏大规模高质量的多层透明数据集,此类生成模型的发展滞后于传统文生图模型。本文通过:(i) 发布首个开源、超高清的PrismLayersPro数据集,包含20万(20万)张带精确透明度掩码的多层透明图像;(ii) 提出无需训练的合成流程,利用现成扩散模型按需生成此类数据;(iii) 推出开源的多层生成模型ART+,其美学表现与现代文生图模型相当。关键技术包括:LayerFLUX用于生成高质量单层透明图像,以及MultiLayerFLUX在人工标注语义布局引导下组合多层输出。通过严格过滤和人工筛选提升质量。在合成数据上微调SOTA模型ART得ART+,在60%的人机对比测试中胜过原版ART,甚至达到FLUX.1-[dev]的视觉水平。本工作为多层透明图像生成奠定了坚实数据基础,助力需要精确、可编辑且视觉出色的分层图像研究与应用。

原文摘要 · Abstract (English)

Generating high-quality, multi-layer transparent images from text prompts can unlock a new level of creative control, allowing users to edit each layer as effortlessly as editing text outputs from LLMs. However, the development of multi-layer generative models lags behind that of conventional text-to-image models due to the absence of a large, high-quality corpus of multi-layer transparent data. In this paper, we address this fundamental challenge by: (i) releasing the first open, ultra-high-fidelity PrismLayers (PrismLayersPro) dataset of 200K (20K) multilayer transparent images with accurate alpha mattes, (ii) introducing a trainingfree synthesis pipeline that generates such data on demand using off-the-shelf diffusion models, and (iii) delivering a strong, open-source multi-layer generation model, ART+, which matches the aesthetics of modern text-to-image generation models. The key technical contributions include: LayerFLUX, which excels at generating high-quality single transparent layers with accurate alpha mattes, and MultiLayerFLUX, which composes multiple LayerFLUX outputs into complete images, guided by human-annotated semantic layout. To ensure higher quality, we apply a rigorous filtering stage to remove artifacts and semantic mismatches, followed by human selection. Fine-tuning the state-of-the-art ART model on our synthetic PrismLayersPro yields ART+, which outperforms the original ART in 60% of head-to-head user study comparisons and even matches the visual quality of images generated by the FLUX.1-[dev] model. We anticipate that our work will establish a solid dataset foundation for the multi-layer transparent image generation task, enabling research and applications that require precise, editable, and visually compelling layered imagery.

图像生成多层图像透明图像数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。