一次生成带真实光影的分层图像,编辑不跑形。
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
- 将光影效果融入前景层,单次前向传播完成生成
- 支持移位、缩放等编辑,无需二次模型推理
- 发布4.8万张带光影的分层图数据集与评测基准
现有图像生成模型在编辑时常导致内容身份失真,如场景偏移或物体外观漂移。分层表示可独立操控元素,但现有方法生成透明前景且缺乏阴影、反射等真实视觉效果,需依赖第二阶段调和模型,反而引入漂移。为此,我们提出LASAGNA,能在单次前向传播中生成具有逼真背景(BG)和含阴影、反射等视觉效果的RGBA前景(FG)。通过将对象相关视觉效果视为前景的一部分,LASAGNA仅用透明度合成即可实现主流用户编辑(如移动、缩放、改色、复制),无需任何后处理模型,彻底避免级联编辑带来的身份漂移。该框架统一处理文本提示、前景、背景及位置掩码等多种条件输入。我们还发布了两个社区资源:LASAGNA-48K——首个公开的4.8万张包含逼真视觉效果的分层图像三元组数据集;LASAGNA-BENCH——首个面向分层生成与编辑的标准化基准,涵盖242个专家标注样本,覆盖六个不同来源。实验表明,LASAGNA在三种生成模式下均优于通用编辑器和先前分层方法,并可在不重新推理模型的前提下支持广泛后编辑操作。
原文摘要 · Abstract (English)
Recent image generation models produce impressive composites, but often fail to preserve the identity of user-provided content when editing specific elements: the surrounding scene may shift, and even the edited object's appearance can drift from the original. Layered representation offer a natural remedy--they allow users to independently manipulate individual elements--but existing layered methods typically produce transparent foregrounds without realistic visual effects such as shadows and reflections, forcing the use of a second harmonization model after every edit, which in turn introduces drift. To overcome these limitations, we present LASAGNA, which generates a photorealistic background (BG) and an RGBA foreground with compelling visual effects in a single forward pass. By treating object-associated visual effects as part of the foreground (FG) layer, LASAGNA supports the dominant class of consumer edits (e.g., translation, scaling, recoloring, duplication) via alpha compositing alone, without invoking any model post-edit, thereby eliminating identity drift introduced by cascade editing pipelines. This single-pass design contrasts with prior layered methods that rely on separate expert models for each task. LASAGNA handles diverse conditional inputs--text prompts, FG, BG, and location masks--within a unified architecture. We further release two community resources: LASAGNA-48K, the first public dataset of 48K layered image triplets with photorealistic visual effects, and LASAGNA-BENCH, the first standardized benchmark for layer-centric generation and editing, comprising 242 expert-annotated samples across six diverse sources. Experiments show that LASAGNA outperforms both general-purpose editors and prior layered methods across three generation modes, and supports a wide range of post-edits without any model re-inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。