让绘画风格与光泽度可独立控制,生成更精准的非写实图像。
Style-Aware Gloss Control for Generative Non-Photorealistic Rendering
- 构建新数据集并训练无监督模型,分离出光泽与风格的层级潜在表示。
- 在多种艺术风格下实现光泽度的精细调控,提升生成图像的可控性。
- 轻量级适配器可接入扩散模型,适合艺术创作与设计领域用户。
人类能从视觉外观推断物体的材质特性,这一能力也适用于艺术作品,其中相似的感知策略引导对绘画或素描的解读。在决定材质外观的因素中,光泽与颜色被广泛认为最重要,且近期研究表明人类可独立于艺术风格感知光泽。为探究光泽与艺术风格在学习模型中的表征方式,我们基于新构建的画家风格物体数据集,训练了一个无监督生成模型,系统性地改变这些因素。分析揭示了光泽与其他外观特征解耦的层级潜在空间,支持对光泽在不同艺术风格下的表示进行详细研究。基于此表示,我们提出一个轻量级适配器,将风格与光泽感知的潜在空间连接至潜在扩散模型,实现对非写实图像中这些因素的细粒度控制。与先前模型相比,我们的方法在学习因子的解耦性和可控性方面均有提升。
原文摘要 · Abstract (English)
Humans can infer material characteristics of objects from their visual appearance, and this ability extends to artistic depictions, where similar perceptual strategies guide the interpretation of paintings or drawings. Among the factors that define material appearance, gloss, along with color, is widely regarded as one of the most important, and recent studies indicate that humans can perceive gloss independently of the artistic style used to depict an object. To investigate how gloss and artistic style are represented in learned models, we train an unsupervised generative model on a newly curated dataset of painterly objects designed to systematically vary such factors. Our analysis reveals a hierarchical latent space in which gloss is disentangled from other appearance factors, allowing for a detailed study of how gloss is represented and varies across artistic styles. Building on this representation, we introduce a lightweight adapter that connects our style- and gloss-aware latent space to a latent-diffusion model, enabling the synthesis of non-photorealistic images with fine-grained control of these factors. We compare our approach with previous models and observe improved disentanglement and controllability of the learned factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。