arXiv:2602.19358cs.CV2026-02被引 3

根据用户提示从一张图中分离出可编辑的图层,实现精准视觉控制。

Referring Layer Decomposition

  • 通过用户提示(如点、框、文字)从单张图像生成完整RGBA图层
  • 构建111万组图像-图层-提示三元组数据集,支持高质量训练与评估
  • 方法零样本泛化能力强,适合图像编辑与生成任务研究者使用

精确、对象感知的视觉内容控制对高级图像编辑和组合生成至关重要。然而,现有方法大多整体处理图像,难以隔离并操作单一场景元素。相比之下,将场景显式分解为物体、环境背景和视觉效果的分层表示,提供了更直观、结构化的理解和编辑框架。为弥合这一差距并实现组合理解与可控编辑,我们提出参考层分解(RLD)任务:从单张RGB图像出发,根据灵活的用户提示(如空间输入、自然语言描述或其组合),预测完整的RGBA图层。核心是RefLade数据集,包含111万组图像-图层-提示三元组,以及10万组人工精修的高保真图层。结合感知一致、符合人类偏好的自动评估协议,RefLade确立了RLD作为可定义且可基准化的研究任务。在此基础上,我们提出RefLayer,一种面向提示条件的图层分解基线模型,在视觉保真度和语义一致性上表现优异,实验表明该方法支持有效训练、可靠评估与高质量分解,并具备强零样本泛化能力。

原文摘要 · Abstract (English)

Precise, object-aware control over visual content is essential for advanced image editing and compositional generation. Yet, most existing approaches operate on entire images holistically, limiting the ability to isolate and manipulate individual scene elements. In contrast, layered representations, where scenes are explicitly separated into objects, environmental context, and visual effects, provide a more intuitive and structured framework for interpreting and editing visual content. To bridge this gap and enable both compositional understanding and controllable editing, we introduce the Referring Layer Decomposition (RLD) task, which predicts complete RGBA layers from a single RGB image, conditioned on flexible user prompts, such as spatial inputs (e.g., points, boxes, masks), natural language descriptions, or combinations thereof. At the core is the RefLade, a large-scale dataset comprising 1.11M image-layer-prompt triplets produced by our scalable data engine, along with 100K manually curated, high-fidelity layers. Coupled with a perceptually grounded, human-preference-aligned automatic evaluation protocol, RefLade establishes RLD as a well-defined and benchmarkable research task. Building on this foundation, we present RefLayer, a simple baseline designed for prompt-conditioned layer decomposition, achieving high visual fidelity and semantic alignment. Extensive experiments show our approach enables effective training, reliable evaluation, and high-quality image decomposition, while exhibiting strong zero-shot generalization capabilities.

图像分割分层表示可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。