用自监督方法学习材料外观的可调控表示,实现精准迁移与编辑。
A Controllable Appearance Representation for Flexible Transfer and Editing
- 基于改进的FactorVAE构建紧凑解耦的外观表征空间。
- 无需人工标注即可有效分离材质与光照属性,支持精细控制。
- 适合作为扩散模型的轻量条件器,适合需要外观编辑的应用场景。
我们提出一种方法,在高度紧凑且解耦的潜在空间中学习材料外观的可解释表示。该表示通过自监督方式,使用改进的FactorVAE在精心设计的无标签数据集上训练,避免了人类标注带来的偏差。尽管没有显式监督,模型仍表现出强解耦性和可解释性,能有效编码材质外观与光照信息。随后,我们将该表示用于训练轻量级IP-Adapter,以条件化扩散流水线,将一张或多张图像的外观迁移到目标几何体上,并允许用户进一步编辑生成结果。由于潜在空间结构良好,用户可在图像空间中直观操控色相、光泽度等属性,实现期望的最终外观。
原文摘要 · Abstract (English)
We present a method that computes an interpretable representation of material appearance within a highly compact, disentangled latent space. This representation is learned in a self-supervised fashion using an adapted FactorVAE. We train our model with a carefully designed unlabeled dataset, avoiding possible biases induced by human-generated labels. Our model demonstrates strong disentanglement and interpretability by effectively encoding material appearance and illumination, despite the absence of explicit supervision. Then, we use our representation as guidance for training a lightweight IP-Adapter to condition a diffusion pipeline that transfers the appearance of one or more images onto a target geometry, and allows the user to further edit the resulting appearance. Our approach offers fine-grained control over the generated results: thanks to the well-structured compact latent space, users can intuitively manipulate attributes such as hue or glossiness in image space to achieve the desired final appearance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。