一歩推論で高品質な多視角素材推定を実現、ノイズ除去が不要
StableIntrinsic: Detail-preserving One-step Diffusion Model for Multi-view Material Estimation
- 采用单步扩散模型,避免多步去噪耗时问题
- 在反照率上PSNR提升9.9%,金属/粗糙度误差分别下降44.4%和60.0%
- 通过细节注入网络保留图像细节,适合高保真材质重建任务
从图像中恢复材质信息是计算机图形学与视觉领域的长期研究课题。近期基于扩散模型的方法虽表现良好,但普遍采用多步去噪策略,推理耗时且与确定性材质估计任务冲突,导致结果方差大。本文提出StableIntrinsic,一种用于多视图材质估计的一步扩散模型,可实现高质量、低方差的材质参数生成。为缓解单步扩散常见的过度平滑问题,该方法在像素空间引入基于材质特性的损失函数;同时设计细节注入网络(Detail Injection Network, DIN),弥补VAE编码带来的细节丢失,进一步提升预测结果的锐度。实验表明,本方法在反照率上达到9.9%的PSNR提升,在金属和粗糙度的均方误差(MSE)上分别降低44.4%和60.0%,显著优于当前最优技术。
原文摘要 · Abstract (English)
Recovering material information from images has been extensively studied in computer graphics and vision. Recent works in material estimation leverage diffusion model showing promising results. However, these diffusion-based methods adopt a multi-step denoising strategy, which is time-consuming for each estimation. Such stochastic inference also conflicts with the deterministic material estimation task, leading to a high variance estimated results. In this paper, we introduce StableIntrinsic, a one-step diffusion model for multi-view material estimation that can produce high-quality material parameters with low variance. To address the overly-smoothing problem in one-step diffusion, StableIntrinsic applies losses in pixel space, with each loss designed based on the properties of the material. Additionally, StableIntrinsic introduces a Detail Injection Network (DIN) to eliminate the detail loss caused by VAE encoding, while further enhancing the sharpness of material prediction results. The experimental results indicate that our method surpasses the current state-of-the-art techniques by achieving a $9.9\%$ improvement in the Peak Signal-to-Noise Ratio (PSNR) of albedo, and by reducing the Mean Square Error (MSE) for metallic and roughness by $44.4\%$ and $60.0\%$, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。