让图像编辑同时保持语义准确和风格统一
Neural Scene Designer: Self-Styled Semantic Image Manipulation
- 用双注意力机制分别控制内容与风格
- 通过对比学习捕捉图像内部一致的风格特征
- 首个针对风格一致性评估的标准化基准
保持风格一致性对图像的连贯性与美学效果至关重要,是有效图像编辑与修复的基本要求。然而,现有方法多聚焦于生成内容的语义控制,常忽视风格一致性这一关键任务。本文提出神经场景设计师(NSD),一种新框架,可在用户指定的场景区域实现照片级真实感修改,同时确保内容语义与用户意图一致,且风格与周围环境保持一致。NSD采用先进扩散模型,引入两个并行交叉注意力机制,分别处理文本与风格信息,以实现语义控制与风格一致性的双重目标。为捕捉细粒度风格表征,提出渐进式自风格表征学习(PSRL)模块,其基于同一图像内各区域风格一致、不同图像间风格相异的直观假设。该模块使用风格对比损失,强化同一图像内表示的相似性,同时强制不同图像间表示的差异性。此外,为解决该任务缺乏标准化评估协议的问题,建立全面基准,包含竞争算法、专设风格相关指标及多样数据集与设置,以促进公平比较。在该基准上的大量实验验证了所提框架的有效性。
原文摘要 · Abstract (English)
Maintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing methods primarily focus on the semantic control of generated content, often neglecting the critical task of preserving this consistency. In this work, we introduce the Neural Scene Designer (NSD), a novel framework that enables photo-realistic manipulation of user-specified scene regions while ensuring both semantic alignment with user intent and stylistic consistency with the surrounding environment. NSD leverages an advanced diffusion model, incorporating two parallel cross-attention mechanisms that separately process text and style information to achieve the dual objectives of semantic control and style consistency. To capture fine-grained style representations, we propose the Progressive Self-style Representational Learning (PSRL) module. This module is predicated on the intuitive premise that different regions within a single image share a consistent style, whereas regions from different images exhibit distinct styles. The PSRL module employs a style contrastive loss that encourages high similarity between representations from the same image while enforcing dissimilarity between those from different images. Furthermore, to address the lack of standardized evaluation protocols for this task, we establish a comprehensive benchmark. This benchmark includes competing algorithms, dedicated style-related metrics, and diverse datasets and settings to facilitate fair comparisons. Extensive experiments conducted on our benchmark demonstrate the effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。