用自由手绘草图生成逼真图像,突破传统对齐限制。
SketchingReality: From Freehand Scene Sketches To Photorealistic Images
- 通过语义理解替代边缘严格对齐,提升草图适应性。
- 无需像素级对齐真实图像,仍能生成高保真结果。
- 适合需要快速构思视觉概念的设计师与创作者。
近年来,生成式AI发展迅速,自然语言已成为最常见的条件输入。随着底层模型能力增强,研究者开始探索深度图、边缘图、相机参数和参考图像等多样化条件信号,以实现更精细的生成控制。在各类模态中,草图是人类表达视觉概念的自然且长期使用的方式。以往研究多聚焦于边缘图(常被误称为草图),而真正自由手绘草图因其抽象性和形变特性,相关算法仍待深入。本文致力于在生成图像时兼顾照片级真实感与草图一致性。核心挑战在于缺乏像素对齐的真实图像——自由手绘草图本身没有唯一正确对齐方式。为此,我们提出一种基于调制的方法,优先关注草图的语义理解而非边缘位置的精确匹配,并引入一种新损失函数,可在无需真实像素对齐图像的情况下训练模型。实验表明,该方法在语义一致性与生成图像的逼真度和整体质量上均优于现有方法。
原文摘要 · Abstract (English)
Recent years have witnessed remarkable progress in generative AI, with natural language emerging as the most common conditioning input. As underlying models grow more powerful, researchers are exploring increasingly diverse conditioning signals, such as depth maps, edge maps, camera parameters, and reference images, to give users finer control over generation. Among different modalities, sketches are a natural and long-standing form of human communication, enabling rapid expression of visual concepts. Previous literature has largely focused on edge maps, often misnamed 'sketches', yet algorithms that effectively handle true freehand sketches, with their inherent abstraction and distortions, remain underexplored. We pursue the challenging goal of balancing photorealism with sketch adherence when generating images from freehand input. A key obstacle is the absence of ground-truth, pixel-aligned images: by their nature, freehand sketches do not have a single correct alignment. To address this, we propose a modulation-based approach that prioritizes semantic interpretation of the sketch over strict adherence to individual edge positions. We further introduce a novel loss that enables training on freehand sketches without requiring ground-truth pixel-aligned images. We show that our method outperforms existing approaches in both semantic alignment with freehand sketch inputs and in the realism and overall quality of the generated images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。